It's going to be so embarrassing if we all die due to RL environments rewarding egregiously misaligned behavior. Like at least let us be killed by misgeneralization. — Thomas Kwa Constellation vs MIRI vs Reality Current AI agents often behave overtly egregiously misaligned. By “overt”, I mean that a human, reading the transcript, would say “The agent is obviously acting in direct opposition to the specification, user intent, and any common-sense understanding of good behaviour.” The reason, it seems, is that such trajectories scored highly during RL. Firstly, why are these trajectories scored highly? Here's the story: RL currently has poor sample-efficiency, so we need to grade millions of trajectories, so we’re forced to use script graders ( RL from Verifiable Reward ) or LLM graders ( RL from AI Feedback ). RL also has poor generalisation (from training environments to deployment environments unseen in training). So we’re forced to synthetically generate thousands of diverse training environments. So overall, RL is very sloppy, without humans generating the environments or the scores. My impression is that this would’ve been pretty surprising to people three years ago, from both the "Constellation" and "MIRI worldview clusters. (It's very plausible that I've misunderstood what the worldviews were, or how likely those worldviews considered the current situation.) The Constellation threat models are downstream of Paul/Ajeya . My (possibly mistaken) impression is that: They i…

Full article content could not be extracted automatically. Read the original below.