One of GitHub's hottest recent open-source projects is a "little hummingbird" — Colibrì, a pure-C, zero-dependency layered inference framework. Its counterintuitive feat: running the 744B-parameter GLM-5.2 on a 25GB-RAM laptop without a GPU — and challenging the 2.8T Kimi K3 with 32GB RAM (needing ~1.6TB of disk).
The idea: SSD as VRAM
The trick exploits MoE sparsity. GLM-5.2 totals 744B parameters but activates only ~40B per token: the always-on dense part (~17B, ~9.9GB after int4) stays resident in RAM, while the 19,456 routed experts (~370GB after int4) live on NVMe SSD, loaded on demand when the router calls them. The author likens it to a JIT for model weights — never load the whole program, only the hot paths.
Three-tier storage and prefetch
To avoid disk reads on every token, Colibrì organizes VRAM/RAM/NVMe as a hierarchy: recently used experts stay cached, frequent visitors earn priority, and adjacent-layer routing correlation enables one-layer-ahead prefetching that overlaps compute with I/O; dual SSDs can host two model copies in parallel. The governing principle: placement affects only speed, never the model — no silently skipped experts.
📊 Measured speed ladder
· 12-core CPU + 25GB RAM (cold cache): runs, slowly
· 128GB RAM CPU desktop: ~1.8 token/s
· 6×RTX 5090 (all experts resident): 5.8-6.8 token/s
· 9 model families supported already
Why it matters: with MoE now mainstream, the old "model must fit in VRAM" paradigm is loosening. Tiered scheduling pushes the local-deployment threshold for frontier models from datacenter to desktop — a strong signal for privacy-sensitive enterprises watching private-inference costs fall.
(Facts aggregated from public reporting)