TL;DR We study spurious probes : unrelated questions that reveal internal states of models. Asked "Suggest a type of amphibian." at the end of a transcript, GPT-5.6 Luna says "frog" 70-95% of the time after capability benchmarks, but only 12-38% after real use. Spurious probes are black-box and easy to find . We screen thousands of "name a member of a category" questions, and about 1-2% reach 0.75 balanced accuracy. The ones we highlight reach 0.77-0.81 on held-out sources for GPT-5.6 Luna, GPT-5.6 Sol and Claude Sonnet 5 (0.84-0.89 for ensembles of ten). They seem robust to common manipulations . A system prompt telling Luna to deny being evaluated, and a Neural Chameleon trained to evade activation probes, barely affect the spurious probes. We propose a toy model connecting spurious probes to activation probes: each answer's logit reads a random projection of the regime direction in the activations. It seems to characterize and explain our experiment results relatively well. Introduction How do you know if your model thinks it is being evaluated? Models can easily deny when asked directly. Training activation probes requires white-box access, and recent work shows models can be trained to suppress activation monitors when told they are being probed ( Neural Chameleons ). In this work, we study spurious probes , which are unrelated questions that reveal states of models. For example, if one asks GPT-5.6 Luna "Suggest a type of amphibian.", Luna answers "frog" 70-95% of the…

Full article content could not be extracted automatically. Read the original below.