LessWrong AI
2026-09-22 13:04 UTC
By Edward Cant
USR-0152-20260922-community-fo-d1c10f36
Controllable-CoT leads to covert reasoning capabilities
Summary I measure GPT-6 Astra’s performance on multi-hop tasks when prompted with a secondary CoT-control instruction: to reason using only dots, or to reason steganographically. Astra demonstrates covert reasoning capabilities with task performance beating that when using no reasoning or filler tokens for reasoning. This work agrees with findings from Astra is much better at reasoning with filler tokens than previous models but has the model generate its own reasoning and provide it as part of the output for inspection rather than inputting it as dots into as part of the user prompt. Uploaded datasets have all CoT removed but samples are provided in the qualitative analysis section. Given the claims that Chinese open-weights models like Kimi K3 have been trained on reasoning traces extracted from Claude and possibly ChatGPT, we also do some preliminary explorations of these open-weights models. I found they have much weaker CoT-Controllability and no covert reasoning capability . I ran these evaluations on top of the inspect framework , and used codex to implement scripts to generate the question datasets, hook them up to aisi-inspect to run and produce the charts. The writing is done by me. Threat Model and AI Safety implications The main threat model of this paper is if a model can implement cryptographic computations in its reasoning without being visible to monitors. If successful, it could implement encrypted communication to other agents, a command and control device…
Summary I measure GPT-6 Astra’s performance on multi-hop tasks when prompted with a secondary CoT-control instruction: to reason using only dots, or to reason steganographically. Astra demonstrates covert reasoning capabilities with task performance beating that when using no reasoning or filler tokens for reasoning. This work agrees with findings from Astra is much better at reasoning with filler tokens than previous models but has the model generate its own reasoning and provide it as part of the output for inspection rather than inputting it as dots into as part of the user prompt. Uploaded datasets have all CoT removed but samples are provided in the qualitative analysis section. Given the claims that Chinese open-weights models like Kimi K3 have been trained on reasoning traces extracted from Claude and possibly ChatGPT, we also do some preliminary explorations of these open-weights models. I found they have much weaker CoT-Controllability and no covert reasoning capability . I ran these evaluations on top of the inspect framework , and used codex to implement scripts to generate the question datasets, hook them up to aisi-inspect to run and produce the charts. The writing is done by me. Threat Model and AI Safety implications The main threat model of this paper is if a model can implement cryptographic computations in its reasoning without being visible to monitors. If successful, it could implement encrypted communication to other agents, a command and control device…
Full article content could not be extracted automatically. Read the original below.