Popular repositories Loading
-
colibri-m3
colibri-m3 PublicPure-C MiniMax-M3 inference engine for CPU-only machines. 6.28 tok/s at 256K context on dual-socket Xeon with 376GB RAM, via mmap zero-copy experts, int2 expert quantization, and persistent scratch…
C 1
-
-
loftlyy
loftlyy PublicForked from preetsuthar17/loftlyy
Brand identity of brands for inspiration.
TypeScript
Repositories
- oxidize Public
Fast, local-first LLM inference in Rust CLI, OpenAI-compatible server, Python bindings, and quantization tools.
- colibri-m3 Public
Pure-C MiniMax-M3 inference engine for CPU-only machines. 6.28 tok/s at 256K context on dual-socket Xeon with 376GB RAM, via mmap zero-copy experts, int2 expert quantization, and persistent scratch buffers.
- Open-Sentry Public
- colibri Public Forked from JustVugg/colibri
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
- snapprune Public
- Luminaweb Public archive
- zapdev Public archive
- miniforge Public
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…