tl;dr: About 80% of released J-lenses target the final layer. On DeepSeek-V3, though, that gives a J-lens dominated by one direction inherited from the final block. It shifts English-vs-Chinese readouts and also inflates one eval. Anthropic's J-lens paper had suggested the final block may specialize in calibrating the next-token prediction. That could make it the block most likely to carry a direction like this, meaning the penultimate layer may be a better default. More generally, this is a case study of how a strong downstream direction can dominate a J-lens and change what earlier layers appear to represent. Overview A J-lens lets you peek inside a model by translating its hidden states into words (a "readout"). It is defined relative to a target layer. Specifically, it asks how a nudge at an earlier layer would change the representation at that target, averaged over many prompts, then reads the result through the model's own unembedding. On DeepSeek-V3, changing the target layer changes what the lens shows you. With the final layer as the target, the J-lens is dominated by a single direction inherited from the last transformer block. That direction shifts the language of the readouts between Chinese and English. [1] This dominant direction arises because in DeepSeek-V3, the final block pushes down all the Chinese tokens when the text is English. This barely changes what the model predicts since those tokens already had almost no probability, but it's a large change to th…

Full article content could not be extracted automatically. Read the original below.