Intro + Background This post is an independent extension of work which I did during Eleuther's SOAR program under Suvajit Majumder's supervision. I found that when given a GCG trigger optimised for output logit entropy, LLMs will randomly take on new personas. This is a new form of prompt injection and could have important safety implications. I recommend reading my SOAR report for context here . In this post, I build on my SOAR work by training linear probes to predict whether an answer will be classified as assistant persona or not. I then investigate by steering with the probes and measuring the balance of personas. Code / data: https://github.com/mild-rgb/CoT-spiking/tree/main/indy_mech_extension / https://huggingface.co/datasets/mild-rgb/indy-mech-extension-qwen3-8b-persona-probes Training probes Method I prefilled Qwen3-8b with full responses (prompt + answer) from the data used in my last post and then do a single forward pass.This reproduces the model's activations when it was generating the tokens without needing to rerun the generation loop. I recorded activations at every even-numbered layer over all of the answer tokens. I then prefilled only the answer and collected activations in the same way as above. I then Z-scored the collected hidden states to account for the first token being an attention sink. This makes answer-only and full-response data comparable. I tried not doing this earlier and got very distorted results. I trained mass mean probes on the Z-scored…

Full article content could not be extracted automatically. Read the original below.