Summary We train a new token—a neologism ( Hewitt et al. )—for a model, but unlike Hewitt et al., we train it on data the model generated while steered with a persona vector. To learn how the model interprets this steering vector, we then ask the model to a) respond in the style of this neologism, and b) explain it. Responses generated with the neologism are substantially more similar to the steering vector (larger projection values) than responses generated with the steering vector itself, while being more coherent and trait-expressive (per an LLM judge). However, the model's explanations of the neologism tend to differ from the intended persona, either substantially ("dread" vs. the intended "evil") or subtly ("warmth" vs. "sycophancy"). Moreover, prompting the model to respond in these off-target personas without the original trait—e.g. "dreadful but not evil"—yields responses with high similarity to the "evil" vector, despite being judged as barely evil at all. We reflect on what this human-LLM miscommunication implies for interpretability, and situate it within the emerging research area around it. Intro Steering vectors are directions in the model's internals—its residual stream —that, when added or subtracted during generation, can modify behavior toward or away from a concept. A large body of work has shown that these vectors have many uses. [1] But how do models interpret their own steering vectors? Presumably, a steering vector for "evil" would be understood by the…

Full article content could not be extracted automatically. Read the original below.