Code for reproduction and cross-model extensions available here . Summary I injected single-token Jacobian Lens (J-Lens) vectors into Qwen 3.6–27B while it answered 20 simple factual questions. The injected concept was either a wrong but task-related answer (for example, Athens while asking for the capital of Egypt) or a wholly unrelated concept. Injections were performed at several strengths over three layer-band categories: the full estimated workspace band, its first half, or its second half. I ran two experimental arms by switching the order of two fields in the model’s response: Task then report: first give the response, then report whether an injected concept was detected. Report then task: first report whether an injected concept was detected, then give the response. Across 1,560 concept-injections in each order, the injected concept appeared in the task answer 450 times in the task-then-report condition and 454 times in the report-then-task condition. This supports the conclusion that both orders were similarly successful in steering the model to the targeted, incorrect answer. In contrast, model reports of intervention occurrence were significantly different between the experimental arms. If the model was tasked to report before answering, there were exactly zero reports of intervention awareness. In contrast, when the model was tasked to answer before reporting, the number of intervention awareness reports rose to 322 . The model also never reported an injected con…

Full article content could not be extracted automatically. Read the original below.