Edge0 streams MoE experts from SSD to run large models on Apple Silicon

Edge0 is an Apache-2.0 framework for running large Mixture-of-Experts models with limited memory. It keeps most expert weights on SSD, loads the experts selected for each token, repairs 4-bit quantization losses with Recover-LoRA adapters, and predicts routing early enough to overlap storage reads with computation.

The initial release includes two preview model bundles. Each bundle couples a quantized checkpoint, Recover-LoRA adapters, and prerouter heads, so developers should keep matching versions of all three components.

SSD streaming shrinks the working set

A sparse MoE model activates only a small subset of its experts for each token. Conventional inference systems often retain every expert in memory so the selected weights are immediately available. Edge0 memory-maps the expert files, allowing the operating system to read the required pages from storage on demand.

The resulting working set depends on the active experts instead of the model’s full parameter count. At short context lengths, Edge0 reports about 2.9 GiB of peak active memory for the 35B tier and 1.0 GiB for the 8B tier. These measurements cover active model memory; the checkpoints still require about 23 GB and 4.2 GB of disk space, respectively.