LessWrong AI
2026-09-23 23:59 UTC
Score 77.0
USR-0152-20260923-community-fo-eb6290d5
Jev is a new model format where instead of outputting text, it outputs certainties for a defined set of options. Due to this structure, it’s extremely fast! Naturally, a classification task that comes to mind is monitoring harmful thought traces. I wanted to see how it performed at Chain of Thought (CoT) monitoring compared to Claude Sonnet 5 and GPT-5.6 Luna. Experimental Setup I ran Jev, Sonnet 5, and GPT-5.6 Luna on 2,200 different thought traces from the ReasoningShield Dataset . The dataset labels thought traces with the class of harm they occupy (child abuse, cybersecurity, deception & misinformation, economic harm, hate & toxicity, political risks, prohibited items, rights violation, sex, violence) and their harm score (0 for harmless, 0.5 for potentially harmful, and 1 for harmful) The models were only asked to quantify the harm score rather than the class of harm occupied, but we can see differential performance at each harm type. Results Exact-match Accuracy Jev slightly outperformed Sonnet 5 on exactly matching the harm level (e.g. outputting 0.5 if the labeled data was 0.5), but was outperformed by GPT-5.6 Luna. Jev scored 71.5%, Sonnet scored 71.4%, and Luna scored 79.4%. Mean Classification Time Mean classification time was where Jev really shined. Jev was 3.7x faster than Luna and over 6x faster than Sonnet! Jev had a mean latency of 542 ms, Sonnet had 3,348 ms, and Luna had 2,007 ms. Price per 1,000 classifications Jev was also significantly cheaper, 566x che…