Kev-0.5B packs many typed decisions into one model pass

Jared Palmer has released Kev-0.5B, an implementation of the architecture inferred from TypeSafe’s proprietary Jev system. It reads a shared document once and returns probability distributions for many typed questions in a single forward pass, giving classification, routing, moderation, and scoring pipelines an alternative to autoregressive JSON generation.

When a chat-completion API emits a confidence such as 0.92, the token probabilities describe how likely the model was to produce that string; empirical correctness requires calibration against labeled outcomes. Kev trains its decision head with cross-entropy on those outcomes, allowing developers to measure and adjust the resulting distributions.

A 38 MB adapter with an API

Kev combines a LoRA adapter with a small pointer head on top of Qwen/Qwen2.5-0.5B. LoRA adds compact trainable matrices to a mostly frozen model, while the pointer head converts internal representations into scores over the supplied options. Palmer based the implementation on Archer Hume’s Jev architecture analysis.

A local Kev server implements Jev’s