SpaceXAI just shipped Grok 4.6, and the headline number is not the benchmark score -- it's the price. The model lands neck-and-neck with Anthropic's Claude Fable 5 on Artificial Analysis's agentic knowledge work benchmark while costing roughly five times less per task to run. That combination is rare enough to be worth paying attention to.

What changed under the hood

Instead of training a new base model, SpaceXAI kept the same 1.5 trillion parameter V9 foundation from Grok 4.5 and concentrated the improvement in post-training: the supervised fine-tuning and reinforcement learning that turn a raw model into a useful one. Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work -- staying with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase, or turning an idea into a polished application.

SpaceXAI says the gains come from a longer supplemental training run, stronger engineering data, and expanded reinforcement learning for coding and knowledge work. The result is a refinement release, not a new architecture -- but the benchmark numbers suggest the post-training investment paid off.

The benchmark that matters here

The key evaluation is AA-Briefcase, Artificial Analysis's agentic knowledge work benchmark. It is worth understanding what this actually tests, because it is quite different from the typical multiple-choice or coding leaderboard.

  • AA-Briefcase evaluates models across four multi-week knowledge work projects, comprising thousands of input files and 91 tasks in total.
  • Across the scenarios, models must complete realistic professional workflows in fields such as data science, product management, and corporate strategy.