LessWrong AI
2026-07-30 19:39 UTC
By Finn Cairns
USR-0152-20260730-community-fo-ef6752b7
Internal State Control is a General Property of LLMs
tl;dr: Lindsey 2025 found models can modulate their internal states : when instructed to “think about” a concept while writing an unrelated sentence, the representation of the concept is more present than when instructed to not think about it. Internal state controllability appears to be a general property of LLMs: the effect replicates in 14 open-weight models from 0.3B to 235B parameters (Qwen3, Gemma 3, Tulu 3) with no clear trend in the think vs. don't-think gap across scale. Since controllability is present even at ≤1B parameters with no size trend, we suspect there is a simpler attention-tagging mechanism at play, rather than metacognition. Current open weight LLMs cannot weaponize this controllability: in a sandbagging setup, the model cannot evade a deception probe when instructed to suppress its signal. Figure 1: Cosine similarity between the concept vector and residual stream at each layer averaged over tokens of the prefilled assistant response, under the think and don’t think prompts, for the Qwen3 model family. The gray region is a baseline of 95% CI of the cosine similarity of unrelated concept vectors, and the shaded bands are ±1 SEM. This replication was done as part of the Second Look Fellowship and supervised by Yixiong Hao and Zephaniah Roe. Our code can be found here . Background Figure 2: The two prompt conditions. Figure design adapted from Lindsey 2025. Activation-based monitoring and interpretability have become important tools for AI oversight. Howev…
tl;dr: Lindsey 2025 found models can modulate their internal states : when instructed to “think about” a concept while writing an unrelated sentence, the representation of the concept is more present than when instructed to not think about it. Internal state controllability appears to be a general property of LLMs: the effect replicates in 14 open-weight models from 0.3B to 235B parameters (Qwen3, Gemma 3, Tulu 3) with no clear trend in the think vs. don't-think gap across scale. Since controllability is present even at ≤1B parameters with no size trend, we suspect there is a simpler attention-tagging mechanism at play, rather than metacognition. Current open weight LLMs cannot weaponize this controllability: in a sandbagging setup, the model cannot evade a deception probe when instructed to suppress its signal. Figure 1: Cosine similarity between the concept vector and residual stream at each layer averaged over tokens of the prefilled assistant response, under the think and don’t think prompts, for the Qwen3 model family. The gray region is a baseline of 95% CI of the cosine similarity of unrelated concept vectors, and the shaded bands are ±1 SEM. This replication was done as part of the Second Look Fellowship and supervised by Yixiong Hao and Zephaniah Roe. Our code can be found here . Background Figure 2: The two prompt conditions. Figure design adapted from Lindsey 2025. Activation-based monitoring and interpretability have become important tools for AI oversight. Howev…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com