Claude Opus 5.5 leads Vals as GPT-6 and MiMo cut costs
Vals AI evaluated six models released in one week through its Vals Index. The weighted aggregate combines finance, coding, and legal benchmarks designed to approximate paid work. Several included agentic tests require models to use tools and complete multistep tasks.
The results put Claude Opus 5.5 at the top, show OpenAI reducing inference costs, and place Xiaomi’s MiMo models near the leading proprietary systems for a fraction of the price. Aggregate rankings still obscure domain regressions, model-routing behavior, and large differences in token use.
Opus 5.5 wins on agents and spends more
Claude Opus 5.5 finished ahead of Claude Fable 5.1 at 68.83%, Claude Opus 5 at 67.21%, and GPT-6 Astra at 66.61%. Vals reports first place on six leaderboards, with published highlights including Terminal-Bench 4.0 at 61.62%, MedScribe at 91.43%, the Vals RSI Index at 37.13%, and a perfect 100% on ProofBench v1.1.
On the Recursive Self-Improvement Index’s language-model training task, Opus 5.5 trained a small model within a 24-hour budget and exceeded the published human baseline. Vals consequently advanced its projected date for full recursive self-improvement by one month from August 2027. That projection is an extrapolation from a constrained benchmark task rather than a direct demonstration of unrestricted self-improvement.
ProgramBench produced one of the clearest gains: Opus 5.5 fully resolved 18.5% of tasks, compared with 7% for Fable 5.1, 5.5% for GPT-6 Astra, and 3% for Opus 5. Its runs finished in 2.4 hours and cost $42.66 per task.
Compared with Opus 5, the newer model raised average cost per test from $18.81 to $22.30. It also regressed on MedCode, Public Benefits, Legal Research, Tax Agent Bench, and HLAB, with losses concentrated in retrieval, source fidelity, and rubric compliance. Its 3.75% score on Harvey’s Legal Agent Benchmark placed it #31 of 64.
Routing and token use complicate deployment
- Anthropic’s server-side safety controls reroute most cybersecurity requests to Claude Opus 4.8. Biology tasks fall back to Claude Opus 5 unless the account is verified. Treating fallback-assisted tasks as failures lowers the reported SRE Bench score from 33.59% to 5.34%.
- Artificial Analysis measured roughly 119,000 output tokens per Intelligence Index task, compared with about 73,000 for Opus 5, 78,000 for Fable 5.1, and 27,000 for GPT-6 Astra. Thinking tokens are billed as output tokens, so always-on adaptive reasoning contributes directly to the higher test cost.
GPT-6 cuts prices as scores hold near prior levels
GPT-6 Sol scored one point below GPT-5.6 Sol’s 63.7%, while reducing cost per test by 47% from $14.21 to $7.56. It improved by 7.3% on Vibe Code Bench, 18.5% on Terminal-Bench Science, and 6.5% on Terminal-Bench 4.0. Legal Research and Tax Agent scores declined.
GPT-6 Luna reduced cost by 46% to $0.42 per test. It posted small gains on Terminal-Bench, MysteryMechanism, and Vibe Code Bench, alongside a small regression on Finance Agent. The two releases shift OpenAI’s price-performance curve through lower inference costs rather than higher aggregate accuracy.
MiMo reaches 59.6% for 20 cents
MiMo V2.6 Flash reached 59.6% on the Vals Index for $0.20 per test, about 1% of the Opus 5.5 cost. It also led CyberBench with a score of 75.4%.
MiMo V2.6 Pro scored 59.5% and rose 27 positions from V2.5 Pro’s #44 finish. Its largest reported benchmark gains were:
- Vibe Code Bench: 51.1%
- ProofBench: 48%
- Legal Research: 31.2%, reaching #11
The Pro release now ranks second among open-weight models at $0.39 per test. Its combination of price, accessible weights, and improved coding scores makes it a practical candidate for throughput-heavy pipelines, subject to testing against the target workload.
Grok 4.7 gains in coding, slips in proofs
Grok 4.7 rose from Grok 4.6’s #19 position to #12. Its Vibe Code Bench score increased by almost 10 points to 86.2%, Terminal-Bench 4.0 gained 11 points, and Harvey’s Legal Agent Benchmark placed it #7. Cost per test increased to about 2.6 times that of Grok 4.6, while ProofBench fell by 25%.
New releases crowd the top 20
Vals reports that 18 of the top 20 models shipped within the past three months. Most recent gains cluster in coding and agentic tool use, with less movement on knowledge and general reasoning tests.
The aggregate gap between the leading proprietary model and the top MiMo result is about 10 percentage points. The corresponding test-cost gap exceeds 100×, creating room for systems that route routine work to cheaper models and reserve premium models for difficult tasks.
A workload-first shortlist
Production selection should pair the Vals shortlist with workload-specific evaluation, including latency, token consumption, fallback routing, retrieval accuracy, and rubric compliance. Those factors can reverse the apparent advantage shown by a single aggregate score.