Google has submitted two models to the ARC Prize verified leaderboard: Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. Both have been officially tested across ARC-AGI-1 and ARC-AGI-2, with results that reveal a sharp cost-performance tradeoff in AI reasoning.

What ARC-AGI actually tests

ARC-AGI measures fluid intelligence: the ability to infer rules from a handful of examples and apply them to a completely novel problem. Each task presents a few input-output grid pairs made of colored squares. The model must identify the transformation rule and apply it to a new input. Answers must match the ground-truth output exactly, no partial credit, and models get at most three attempts per task.

The benchmark is deliberately resistant to memorization. Most benchmarks reward knowledge; ARC-AGI-2 rewards adaptability. Average individual human performance sits at 66% on ARC-AGI-1, the human panel completion rate is 100%, and the grand prize threshold is above 85% on ARC-AGI-2.

Flash vs. Flash-Lite: the numbers