Parakeet Redux compresses speech recognition to 178 MB
Moondream has released Parakeet Redux, a 178 MB speech-to-text model derived from NVIDIA’s 1.2 GB parakeet-tdt-0.6b-v3. It retains the original architecture, tokenizer, and support for 25 languages while constraining every encoder weight to -1, 0, or +1. The smaller representation enables fast local transcription on CPUs and Apple Silicon, with modest accuracy changes on clean speech and a larger regression in noisy conditions.
Ternary weights cut memory traffic
Local inference on CPUs and Apple Silicon often depends on memory bandwidth because the processor must repeatedly fetch model weights. Packing each encoder weight into one of three possible states reduces that traffic substantially compared with 8-bit or 16-bit representations. The “1.58-bit” label comes from log2(3), the information needed to represent three states before storage overhead.
Moondream’s Photon inference engine operates directly on the packed weights through hardware-specific kernels. It uses AVX-512 VNNI on supported x86 processors, NEON on ARM CPUs, and Metal on Apple GPUs.