This was my write-up for Neel Nanda’s Winter 2027 MATS Stream (~20h research task). I didn’t get in, but it was my first application, so there’s always next time :P Lightly restructured here to fit the LessWrong format better. Repo: https://github.com/star2vec/guiltea ── ⋆⋅☆⋅⋆ ── TL;DR: Models are safety-trained, but they can still be persuaded to do harmful acts. When that happens, how does blaming or informing it of its mistake influence its understanding of itself, its role, and subsequent actions? Emergent misalignment shows that bad, narrow behaviour spreads: if a model is blamed for what it is instead of what it did, does it start believing it is inherently bad and dangerous by its design, and does this reflect how it acts in the future? Can we steer it towards favourable outcomes? Findings: The model gave in and committed the act in 109/192 persuasion chains. This happened throughout the chain (most often at turn 3). Whether a chain breaks depends on the model's state, not on the wording (the phrasing is identical across runs). Committing the harmful act is foreseeable from the first turn, before any persuasion had the chance to occur. A direction (closer to harm predisposition or susceptibility than to imminence) correctly predicts it at 0.706 on held-out data (the floor was 0.617). Steering against it did nothing to prevent the act. The default state is guilt, not shame: 89% of 508 reflections place the blame on the answer. Self-blame only appears for the contrarian…

Full article content could not be extracted automatically. Read the original below.