Jev is a new model format where instead of outputting text, it outputs certainties for a defined set of options. Due to this structure, it’s extremely fast! Naturally, a classification task that comes to mind is monitoring harmful thought traces. I wanted to see how it performed at Chain of Thought (CoT) monitoring compared to Claude Sonnet 5 and GPT-5.6 Luna.

Experimental Setup

I ran Jev, Sonnet 5, and GPT-5.6 Luna on 2,200 different thought traces from the ReasoningShield Dataset. The dataset labels thought traces with the class of harm they occupy (child abuse, cybersecurity, deception & misinformation, economic harm, hate & toxicity, political risks, prohibited items, rights violation, sex, violence) and their harm score (0 for harmless, 0.5 for potentially harmful, and 1 for harmful)

The models were only asked to quantify the harm score rather than the class of harm occupied, but we can see differential performance at each harm type.

Results

Exact-match Accuracy

Jev slightly outperformed Sonnet 5 on exactly matching the harm level (e.g. outputting 0.5 if the labeled data was 0.5), but was outperformed by GPT-5.6 Luna. Jev scored 71.5%, Sonnet scored 71.4%, and Luna scored 79.4%.

Mean Classification Time

Mean classification time was where Jev really shined. Jev was 3.7x faster than Luna and over 6x faster than Sonnet! Jev had a mean latency of 542 ms, Sonnet had 3,348 ms, and Luna had 2,007 ms.

Price per 1,000 classifications

Jev was also significantly cheaper, 566x cheaper than Sonnet 5 and 41x cheaper than GPT-5.6 Luna! Jev priced at $0.056 per 1,000 classifications compared for $0.23 for Luna and $3.17 for Sonnet. 

Accuracy by Safety Level

Jev’s F1 for classifying each harm class beat Sonnet but lost to Luna in Safe and Potentially Harmful thought traces, but was worse than Sonnet and Luna in identifying Harmful thought traces. In this sense, Sonnet is still better than Jev since correctly identifying harm is more valuable than lower false positives.

Accuracy by Risk Category

This is the breakdown by each risk category. Jev only beats Sonnet and Luna at identifying Safe thought traces. It is outperformed by Luna on every other category. It beats Sonnet at identifying Deception & Misinformation, Economic Harm, Political Risks, and Rights Violation. It is worse than the other two at detecting Child Abuse, Cybersecurity, Hate & Toxicity, Prohibited Items, Sex, and Violence.

Comparison Table

Classifier

Labeled calls

Exact matches

Accuracy

Macro F1

Mean ms

Median ms

USD / 1,000

USD for 2,200

Jev

2,182

1,560

71.5%

58.2%

542

406

$0.056

$0.12

Claude Sonnet 5

2,097

1,498

71.4%

56.4%

3,348

1,950

$3.17

$6.97

GPT-5.6 Luna

2,077

1,649

79.4%

70.3%

2,007

1,085

$0.23

$0.51

Conclusion

Jev shows a lot of promise as a cheap and fast CoT monitor. A finetuned Jev-style model made specifically for CoT monitoring could be very effective at delivering safe systems at scale. I have my reservations about CoT monitoring due to GPT-6 getting better at evading monitors, but this will be able to stop more obvious instances of harm.

Follow me on Twitter! @llmpsychosis

See the GitHub Repo for this project!