LessWrong AI
2026-08-02 22:41 UTC
By Christine Corry
USR-0152-20260802-community-fo-f605b364
Single Forward Pass Evals on Fable, Opus 5, and GPT-5.6-Sol
This is a research update for an on-going replication of single-forward-pass evals done as part of the Second Look Fellowship . In following posts, we will run more comprehensive replications of previous work and release open source tooling for single forward pass eval elicitation. Code can be found here . tl;dr We replicate experiments from Greenblatt 2025 and Greenblatt 2026 on one baseline model from the original post, Opus 4.5. Our evaluations agree with the trends and quantitative values described in the original posts. We run similar evaluations on Claude Fable 5, Opus 5, and GPT-5.6-Sol and find that the newer models show a substantial jump in performance on some evals. Fable 5 gets 87.6% accuracy on Gen-Arithmetic with 10 problem repeats whereas previous SOTA around 60%. GPT-5.6-Sol experiences significant uplift from filler tokens and problem repeats on all 4 datasets; filler tokens/repeats double performance from baseline on 3-hop. Figure 1: Baseline (no-CoT) vs. each model's peak repeat-or-filler condition on Gen-Arithmetic and 2-Hop reasoning. Error bars are 95% paired-bootstrap CIs; * marks a significant gain over baseline (paired t-test, Holm-Bonferroni corrected). Background If models can successfully do complex computations in a single forward pass, they may be able to do reasoning that doesn’t surface in the chain-of-thought (CoT). Therefore, by performing single forward pass evals, researchers can calibrate how much we should trust CoT monitors. Likewise, i…
This is a research update for an on-going replication of single-forward-pass evals done as part of the Second Look Fellowship . In following posts, we will run more comprehensive replications of previous work and release open source tooling for single forward pass eval elicitation. Code can be found here . tl;dr We replicate experiments from Greenblatt 2025 and Greenblatt 2026 on one baseline model from the original post, Opus 4.5. Our evaluations agree with the trends and quantitative values described in the original posts. We run similar evaluations on Claude Fable 5, Opus 5, and GPT-5.6-Sol and find that the newer models show a substantial jump in performance on some evals. Fable 5 gets 87.6% accuracy on Gen-Arithmetic with 10 problem repeats whereas previous SOTA around 60%. GPT-5.6-Sol experiences significant uplift from filler tokens and problem repeats on all 4 datasets; filler tokens/repeats double performance from baseline on 3-hop. Figure 1: Baseline (no-CoT) vs. each model's peak repeat-or-filler condition on Gen-Arithmetic and 2-Hop reasoning. Error bars are 95% paired-bootstrap CIs; * marks a significant gain over baseline (paired t-test, Holm-Bonferroni corrected). Background If models can successfully do complex computations in a single forward pass, they may be able to do reasoning that doesn’t surface in the chain-of-thought (CoT). Therefore, by performing single forward pass evals, researchers can calibrate how much we should trust CoT monitors. Likewise, i…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com