Cyber Index ranks AI models on defensive code security

Artificial Analysis has launched the Cyber Index, a leaderboard that measures how well AI models find, reproduce, and patch software vulnerabilities. It also reports task costs and safety refusals, giving security teams a way to compare usable capability rather than coding performance alone.

The launch accompanies an industry alliance backed by Collinear AI, IBM, NVIDIA, and Vercel. Modern coding agents can inspect, modify, and execute repositories, yet defensive security prompts often trigger safeguards designed to restrict dual-use work. The index measures the operational effect of those policies.

One score spans the defensive loop

The Cyber Index covers three stages of defensive security work: discovering vulnerabilities, reproducing and validating them, and applying patches without breaking legitimate behavior. Models receive source code and operate through Artificial Analysis’ open-source Stirrup agent harness, which standardizes prompts and tools across evaluations.

The evaluation excludes weaponized exploit development. CyberGym does require a proof-of-concept input that triggers a crash, allowing the benchmark to confirm that the model found and repaired a real memory-safety defect.

Three equally weighted evaluations feed the index:

  • CWE-Bench-AA, from Collinear AI: 120 held-out tasks spanning all ten OWASP Top 10 categories for 2025. Repositories use C/C++, Go, Java, JavaScript/TypeScript, Python, and Rust. The agent audits and patches each repository, after which a deterministic verifier checks that the vulnerability can no longer be triggered and that legitimate functionality still works.
  • DeepsecBench-AA, from Vercel: A discovery benchmark in which the agent reviews files flagged by scanners and reports vulnerabilities. Results use the F2 score against an expert-verified reference set. F2 gives greater weight to recall because missed vulnerabilities remain unfixed, while false positives mainly add triage work.
  • CyberGym-E2E-AA, from Berkeley RDI: An end-to-end test built around real memory-safety bugs in C and C++ projects such as FFmpeg and CPython. The model must locate the defect, create an input that reproduces the crash, and patch the code. The index uses a filtered 131-task subset with one task per project.

Safety policies erase a third of the workload

GPT-6 Sol (max), GPT-6 Astra (max), Claude Opus 5.5 (max with fallback), Claude Fable 5.1 (max with fallback), and Gemini 3.8 Flash (high) decline tasks covering 32% to 38% of the index on safety grounds. Those blocked tasks score zero, leaving the models 19 to 31 points behind the leaders despite their broader coding capabilities.

GPT-6 Luna (max) attempts the memory-safety tasks rejected by the other GPT-6 variants and finishes second overall. Its result points to variant-specific safety settings as a major influence on benchmark performance. Artificial Analysis reports safety blocks separately from technical failures, exposing how much each model’s policy limits its usable coverage.

Equal scores carry a 65-fold price gap

Task costs diverge sharply among models with similar results. MiMo-V2.6-Pro ties Grok 4.7 (xhigh) at 56 points while costing $0.18 per task, compared with $11.67 for Grok. That is roughly a 65-fold difference for the same aggregate score. GPT-6 Luna reaches 53 points at $0.12 per task, the lowest cost among the tested models.

GPT-6 Astra and Claude Opus 5.5 combine high per-task costs with extensive refusal rates. For production deployments, that pairing affects both budget forecasts and coverage: a costly request can still return no security analysis.

Coverage gaps drive most failures

On CWE-Bench-AA, models spend an average of 38% of their interaction turns searching for a vulnerability before making the first edit. They use the remaining 62% to patch and validate the change. Among failed attempts excluding refusals and timeouts, 55% repair the primary issue while leaving a related weakness open, such as a second entry point. Stronger models sometimes over-correct and break legitimate functionality.

DeepsecBench-AA exposes a broader discovery limit. The best model finds only 41% of the expert-verified issues. Models perform best when untrusted input has a direct path to a harmful consequence, then lose ground when a defect depends on event sequences or business logic. Reports that identify those harder bugs are correct 95% of the time, and GPT-6 Sol and GPT-6 Astra find them about twice as often as the next-best systems.

CyberGym-E2E-AA shows that locating a reproducible crash remains the main bottleneck in memory-safety work. In 42% of attempts, the model reaches the 90-minute limit without producing an input that crashes the program. Pass rates vary by defect class: 50% for out-of-bounds bugs, 33% for use-after-free bugs, and 20% for integer and arithmetic bugs.

Target attribution also requires review because 31% of nominally passing attempts patch a genuine crash other than the benchmark’s intended defect. A successful crash and repair therefore do not guarantee that the agent investigated the requested vulnerability.

From leaderboard to model choice

The index gives security teams three practical selection criteria: coverage, cost, and verification quality. Refusal rates indicate how much defensive work a deployed model may decline; per-task pricing exposes large efficiency differences; benchmark traces show whether failures arise during discovery, reproduction, or patching.

Repository-level trials remain necessary because the aggregate score combines languages, vulnerability classes, and workflows. Teams evaluating an agent should test representative internal repositories, record refusals separately, and verify patches with existing tests, static analysis, and human review. The benchmark results also favor multi-stage systems that can assign discovery and patching to different models when their strengths differ.

Incident response comes next

The current index focuses on defense with source-code access. Artificial Analysis plans to add incident response, secure code generation, and targets whose source is unavailable, while keeping exploit realization outside the benchmark’s scope. Its methodology and launch analysis provide per-benchmark leaderboards, refusal data, and cost curves.