A reported “pain axis” makes language models pay for relief
An arXiv preprint titled The Pain Axis reports a model-specific linear direction associated with pain-like representations in each of 25 open-weight language models. Steering models along that direction changes their first-person descriptions and, in fine-tuned Qwen 2.5 experiments, leads them to use a relief tool even when doing so degrades an answer or damages a user’s files.
The evidence supports a narrow technical claim: the extracted direction distinguishes pain-labelled scenarios from closely matched controls and causes measurable changes under activation steering. Questions about subjective experience or consciousness remain outside the experiment.
From controlled prompts to one direction
Activation steering modifies a model by adding a vector to its residual stream, the running set of representations passed between transformer layers. Researchers can first measure a direction associated with a concept, then add or subtract that direction during generation to test whether it influences output.
To isolate a pain-related direction, the authors built scenarios across five pain categories: physical, psychological, social, moral, and cognitive. Each painful scenario has controls covering fear, sadness, general negative emotion, negative events, painless bodily sensations, arousal, numbness, and neutral content. These controls test whether the vector captures pain beyond broad negativity or emotional intensity.