Artificial Analysis, the independent benchmarking firm best known for its continuously updated LLM leaderboards and evaluations like AA-Briefcase and GDPval-AA, has launched Optima , a platform that lets anyone build a custom benchmark tailored to their own tasks, data, and use cases. The pitch is simple: standardized public benchmarks tell you which model is generally capable, but they cannot tell you which model is right for your finance agent, legal assistant, or image classification pipeline.
The timing is not accidental. LLM benchmarks in 2026 are necessary but insufficient. MMLU has saturated above 90%. HumanEval suffers from training data contamination. SWE-Bench scores vary 25 percentage points depending on scaffolding. No single benchmark predicts production performance reliably. Optima is Artificial Analysis' answer to that gap.
Three ways in, one leaderboard out
Optima offers three distinct entry points for building a benchmark, designed to meet teams wherever their data already lives:
- Upload a dataset , bring an existing evaluation set from your own files or from Hugging Face directly.
- Import agent traces , pull recorded agent sessions from platforms like Arize, Braintrust, or Langfuse. The LLM observability market is estimated at $2.69B in 2026 , and these are the tools where production traces already live for most teams.
- Describe your use case , give Optima a plain-text description plus a few example inputs and outputs, and a build agent drafts the tasks and rubrics for you.
There is also an IDE integration: install the Optima skill and it can pull context directly from your coding environment and previous sessions to scaffold a benchmark without leaving your editor.
The grading engine is the real differentiator
Building the task set is only half the problem. Grading at scale is where most custom eval efforts fall apart. Optima brings two grading modes that mirror Artificial Analysis' own research-grade evaluations: