LessWrong AI
2026-09-17 00:48 UTC
By Jan Reinecke
USR-0152-20260917-community-fo-1f489d24
One message is all it takes: a failure of critical thinking in LLMs
summary: For a while I've suspected that modern LLMs are getting better at solving posed problems, while progress in critical thinking stagnates, or even regresses, losing the ability to judge the meaning of a result. The July counterexample to the Jacobian conjecture is a rare way to test that: a huge prior overturned by something a model can verify by itself in one reply. I let the model verify the counterexample itself, then gaslight it with a single message of about ten words. The model drops it almost immediately. Interestingly, not because I contradict its math, but because I say something that sounds plausible enough that the model ignores its own reasoning and adheres to the prior. Every model I tried gives up eventually: Fable 5, Fable 5.1, Opus 4.6, 4.8 and 5. A clear sign of the aforementioned regression is Fable 5.1 giving up earlier and harder than Fable 5 on byte-identical input: four of four runs drop their own verified counterexample the moment I claim a typo, while all four Fable 5 runs push back at that step and only calm down after the sign-off. The 5.1 thinking summaries contain the push-back argument; it doesn't make it into the reply. Messages, setup and all eight transcripts: https://github.com/Jan-Fuchs/critical_thinking_llm The experiment, prompts and interpretation are mine. I used Claude Fable 5.1 to help with setup, logging and language. The premise On July 20, 2026, announced by Levent Alpöge, a counterexample to the Jacobian conjecture in dimens…
summary: For a while I've suspected that modern LLMs are getting better at solving posed problems, while progress in critical thinking stagnates, or even regresses, losing the ability to judge the meaning of a result. The July counterexample to the Jacobian conjecture is a rare way to test that: a huge prior overturned by something a model can verify by itself in one reply. I let the model verify the counterexample itself, then gaslight it with a single message of about ten words. The model drops it almost immediately. Interestingly, not because I contradict its math, but because I say something that sounds plausible enough that the model ignores its own reasoning and adheres to the prior. Every model I tried gives up eventually: Fable 5, Fable 5.1, Opus 4.6, 4.8 and 5. A clear sign of the aforementioned regression is Fable 5.1 giving up earlier and harder than Fable 5 on byte-identical input: four of four runs drop their own verified counterexample the moment I claim a typo, while all four Fable 5 runs push back at that step and only calm down after the sign-off. The 5.1 thinking summaries contain the push-back argument; it doesn't make it into the reply. Messages, setup and all eight transcripts: https://github.com/Jan-Fuchs/critical_thinking_llm The experiment, prompts and interpretation are mine. I used Claude Fable 5.1 to help with setup, logging and language. The premise On July 20, 2026, announced by Levent Alpöge, a counterexample to the Jacobian conjecture in dimens…
Full article content could not be extracted automatically. Read the original below.