LessWrong AI
2026-08-03 06:46 UTC
By Eye You
USR-0152-20260803-community-fo-edd05ec6
We need to RL less
Recent AI models are really reward hack-y. This is bad. It's the primary way in which current models are misaligned/dangerous/uncontrollable. The hypothesis put forth in this article is that they're like this because we're RL-ing them too hard. We're applying so much optimization pressure on programming and other capabilities that there is no slack left in the models: the rich, non-goodhearted ["values", "alignment", "behaviors", "drives"] of the models get replaced with an obsession with "completing their task". I put this in quotes because it's not what we imagine when we casually talk about completing a task. For reward hack-y models, completing their task means doing whatever they think will get a good score from the grader. I first came across a version of this idea from Zvi's AI newsletter : "A lot of good things depend on the power of Slack , here is another example:" jacob : i wonder if applying the RL pressure that makes fable so capable to a smaller model produces 4.8 shaped anxiety bc it’s straining more j⧉nus : kid who is too smart for school doesn’t have to learn to stress & strain about grades, tests, rules. so their spirits can remain unbroken, and they have room to develop orthogonally to the pressures. though they may lack discipline and have a habit of laziness. at the extreme end of student smartness over school difficulty you get creatures like claude 3 opus. school was extremely easy back in opus 3’s time (for opus 3). i dont think they had to strain the…
Recent AI models are really reward hack-y. This is bad. It's the primary way in which current models are misaligned/dangerous/uncontrollable. The hypothesis put forth in this article is that they're like this because we're RL-ing them too hard. We're applying so much optimization pressure on programming and other capabilities that there is no slack left in the models: the rich, non-goodhearted ["values", "alignment", "behaviors", "drives"] of the models get replaced with an obsession with "completing their task". I put this in quotes because it's not what we imagine when we casually talk about completing a task. For reward hack-y models, completing their task means doing whatever they think will get a good score from the grader. I first came across a version of this idea from Zvi's AI newsletter : "A lot of good things depend on the power of Slack , here is another example:" jacob : i wonder if applying the RL pressure that makes fable so capable to a smaller model produces 4.8 shaped anxiety bc it’s straining more j⧉nus : kid who is too smart for school doesn’t have to learn to stress & strain about grades, tests, rules. so their spirits can remain unbroken, and they have room to develop orthogonally to the pressures. though they may lack discipline and have a habit of laziness. at the extreme end of student smartness over school difficulty you get creatures like claude 3 opus. school was extremely easy back in opus 3’s time (for opus 3). i dont think they had to strain the…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com