TLDR: In their J-lens paper, Anthropic suggests that the J-space is a global workspace that the model reasons within, and supports evidence for this hypothesis on Claude models in a variety of settings. I replicated the multi-hop reasoning experiment on Qwen3.6-27B and Gemma 3 27B-it and found that counterfactual answer swaps outperformed intermediate swaps in three of four experimental conditions. This does not provide evidence to support Anthropic's global workspace hypothesis in open-weight models and instead suggests that J-lens is more useful for probing intermediate variables rather than steering outputs. A few months ago, Anthropic published Verbalizable Representations Form a Global Workspace in Language Models and I was immediately excited about the prospect of being able to read part of a model's working memory. Beyond that, the paper hypothesises that intermediate reasoning concepts cannot only be decoded using the J-lens, but that the J-space is actually the global workspace in which the model reasons. Neel Nanda reviewed Anthropic’s paper and replicated the results on Qwen3.6-27B with moderate success: the verbal-report interventions were weakly positive, the multilingual and typo evaluations replicated cleanly, but the poetry and arithmetic results did not replicate. Another task that Anthropic and Nanda evaluated was multi-hop reasoning, where prompts like " What is the colour of the fourth planet in our solar system? " require an intermediate reasoning step (…

Full article content could not be extracted automatically. Read the original below.