This is a linkpost for the write-up on my site ; the full body is below, and the code, decisions ledger and devlog are in the repo . Base-model z-score calibration of the J-lens helps elicit hidden secret words from Cywiński et al.'s taboo organisms: 0.805 leave-one-out accuracy against 0.665 for their protocol on Gemma-2-9B-it, and the only non-zero readout on Qwen3-1.7B. The J-vs-logit part of that gap is a point estimate at n = 20 (paired sign-flip p ≈ 0.19; p ≈ 0.23 against a z-scored logit lens), so "calibration helps" is the finding and "J-lens beats logit lens" is suggestive. Credit: phoenix's comment on the workspace post claimed the final-layer bias correlates with log token frequency at r ≈ 0.67. I measure r = 0.70 on GPT-2, so the GPT-2 cell below partly confirms it. Executive summary Problem. I want to quantify the part of the J-lens readout that is not affected by an activation's meaning, which I refer to as the "non-context offset". If the lens is going to be used to read hidden content, this offset is where the lens will fail. I found that the offset is mostly the model's own token-frequency, so subtracting it removes useful information. But scaling with variance helps it. The terminology I use: the logit lens is norm and unembed applied to a residual activation at layer L. The J-lens (Anthropic's global workspace paper , discussed on LessWrong ) is similar, but passes the activation through a fitted Jacobian first. The R-lens (the R-lens post by camilablank,…

Full article content could not be extracted automatically. Read the original below.