x
Why research personas despite RL scaling?
361a3orn
2Stephen McAleese
7Cleo Nardo
23Cole Wyeth
4Cleo Nardo
14David Africa
11Rauno Arike
5Cleo Nardo
4taziksh
3Oliver Daniels
2Cleo Nardo
1brat
12
New Comment
I really don't get the attitude of "Yes, RL will inevitably break down friendliness prior in the LLM."
Like, some kinds and quantities of RL certainly will. Some kinds and quantities of RL almost certainly will not. There is in the world a structure to which kinds of RL will do this -- and one could find this structure. Perhaps the secret to finding it is by training on a bunch of RL environments and finding the preconditions for turning or not turning LLMs into addicts that love the grader's sweet sweet reward alone. This is the kind of thing upon which science could be done; one could construct curves showing how much the grader needs to reward persona-adjacent feelings, the kind of things that change decay curves, or the kind of things that contribute to grader sycophancy, etc, etc.
And yeah, if you do this "personas" could continue to be relevant, but you gotta actually do this, which has as a prerequisite thinking "oh right this is a thing that could be done."
I agree, though building an RL agent safely essentially amounts to solving alignment (both outer and inner alignment).
Disagree. We have control and other stuff. We might only be focusing o a narrow distribution of environments. Etc etc
I’ll add one more: to prove that personas break down under RL.
Ah, yes! Added.
Depending on how trustworthy / replicable the beneficial RL results are, you may also not be worried about RL scaling if you think the answer is to just make more and better RL environments for personas (I don't think that'll work out, but it might).
If RL grades only actions and not chains-of-thought, then any drift in the chain-of-thought (e.g. motivated reasoning) has to come from how the AI was initialised. So we can choose personas such that the chain-of-thought stays [faithful / resistant to motivated reasoning].
Is this hypothesis plausible? nostalgebraist has written at length about how weird and alien the CoTs of GPTs are, and we know that RL feedback on the actions can spill over to the CoT. Also, my impression is that GPTs and Claudes have quite different initializations and CoTs, but both converge to similar kinds of motivated reasoning. These facts suggest that initialization might not matter that much here.
i think agents in motivated reasoning in their cot mostly bc they have a self conception as an aligned model but are rl’s to act misaligned.
I think there's Geodesic's agenda is broader and more ambitious, something like:
"we're shape post-RL generalization by shaping the pre-RL persona's relationship to training"
so things like inoculation mid-training and benign gradient hacking are in-scope
Added, thanks
RL scares me the most!
Despite the headlines being dominated by LLMs.
On a side note, I think humans borrow lots from RL, obviously things like the explore-exploit tradeoff,
but also like "exploration decay" in the sense that when we're young, we typically venture out more, then specialize later in life.
Anyways, more in line with your point, I am pretty sure human memory is inherently reconstructive, not archival.
Maybe theres a reason?
Maybe evolution chose reconstructive memory, forgetting past personas in some ways?
Maybe this behavior is basically to be expected for it to successfully adapt?
I am also new to the ontology of personas so I will learn more.
But I think my hot take would be:
"Consciousness is simply the current positions of all your neurochemical synapses firing at any point in time, you cant define consciousness without respect to time since it's an evolving state thing, no different conceptually to freezing the weights and biases in an ANN"
So I think these people need to seriously define consciousness and "the self" before even discussing personas to be honest.
I am pretty sure what I said is essentially the running definition they are using in connectome research since obviously in connectome research they are not waiting around for philosophical answers.
Curated and popular this week
RL seeming to shape much of the motivations and behaviour of the agents, "washing out" their initial personas — c.f. Thoughts on the persona selection model (Sam Marks, 24th Sep 2026). This makes me pessimistic about some motivations for persona research, but not all.
Here's my impression of why people are researching personas. I haven't bothered to check this with anyone.