Artificial Analysis has patched its Intelligence Index from v4.1 to v4.1.1. The scope is narrow by design: no new benchmarks, no reshuffled weights. The fixes target grading errors and grader inconsistency across evaluations. Claude Opus 5 holds the top spot with an Index score of 63.

What changed

Two categories of fixes landed in this patch. The τ³-Banking benchmark was updated to its v1.0.1 release from Sierra Platform, and three evaluations received new grader models.

The v1.0.1 grading update corrects errors in the banking_knowledge task, specifically in trajectories where an agent recovers from a mistake mid-task. These "unhappy paths" were previously scored incorrectly. Scores from earlier versions are not directly comparable to v1.0.1 results.