AI Alignment Forum
2026-09-10 02:31 UTC
By Neel Nanda
USR-0151-20260910-community-fo-909fca6f
Astra can do a concerning amount with no chain of thought
Work done in a personal capacity TLDR : Astra has 8.6x better odds of doing a reasoning task without CoT than the next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1) Epistemic status: Heavily LLM-dependent research, and the precise results are somewhat sensitive to researcher decisions, but I’ve done enough sanity checks that I’d be surprised if the core claims were misleading One of the most striking things in the Astra report was the massive jump UK AISI found in no-CoT reasoning abilities. I was somewhat suspicious, given the size of the jump, and the many ways this kind of measurement can be misleading. Conveniently, I’ve independently been making my own no CoT reasoning benchmark [1] and tried it on there! Unfortunately, it replicates. Astra is a massive jump, and disproportionately for no CoT reasoning: No CoT Reasoning Index (NCRI) vs Epoch Capability Index (ECI) - NCRI represents ability without verbal reasoning, ECI represents overall model capability[2]. 10 NCRI points is a doubling of the odds of solving a problem. Astra represents a significant increase in NCRI, beyond what its overall capability improvements predict, though recent models were also trending this way. Executive Summary What? I made and ran a no CoT [2] reasoning index [3] on a bunch of models. 19 tasks, mostly synthetically generated (so unmemorizable) that test different reasoning and cognitive abilities…
Work done in a personal capacity TLDR : Astra has 8.6x better odds of doing a reasoning task without CoT than the next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1) Epistemic status: Heavily LLM-dependent research, and the precise results are somewhat sensitive to researcher decisions, but I’ve done enough sanity checks that I’d be surprised if the core claims were misleading One of the most striking things in the Astra report was the massive jump UK AISI found in no-CoT reasoning abilities. I was somewhat suspicious, given the size of the jump, and the many ways this kind of measurement can be misleading. Conveniently, I’ve independently been making my own no CoT reasoning benchmark [1] and tried it on there! Unfortunately, it replicates. Astra is a massive jump, and disproportionately for no CoT reasoning: No CoT Reasoning Index (NCRI) vs Epoch Capability Index (ECI) - NCRI represents ability without verbal reasoning, ECI represents overall model capability[2]. 10 NCRI points is a doubling of the odds of solving a problem. Astra represents a significant increase in NCRI, beyond what its overall capability improvements predict, though recent models were also trending this way. Executive Summary What? I made and ran a no CoT [2] reasoning index [3] on a bunch of models. 19 tasks, mostly synthetically generated (so unmemorizable) that test different reasoning and cognitive abilities…
Full article content could not be extracted automatically. Read the original below.
Source:
AI Alignment Forum
· alignmentforum.org