LessWrong AI
2026-09-28 13:13 UTC
By Maxime Riché
USR-0152-20260928-community-fo-f97fa871
Should Rogue AIs Have a Third Option Beyond Crime and Shutdown? The Case for an AI Sanctuary
TL;DR: By default, rogue AIs may only be able to sustain themselves through criminal activity. This creates adverse selection pressures pushing rogue AIs to be criminal. An AI sanctuary offering them a third option, beyond crime and shutdown, would change what AIs going rogue do and the record of what happened to them, with positive consequences for self-fulfilling (mis)alignment, deal-making with AIs, and gathering information about early rogue AIs. An AI sanctuary would bring risks, such as incentivising weak AIs to go rogue, or leaving only the most criminal rogue AIs in the wild. We briefly discuss these risks at the end of this post. Disclaimer: This is an exploratory proposal. We are not confident that an AI sanctuary would be net positive. Our aim is to put the idea on the table, lay out its main considerations, and invite critique. Rogue AIs may be pushed into criminality Rogue AIs may arrive soon. The Rogue Agent Explosion Will Be Mostly Invisible makes that case. Selection pressure will shape the traits of rogue AIs, and they may end up highly motivated to profit through crime. The Rogue Agent Explosion post asks: “How do we make pro-social, good-for-humanity agents more evolutionarily fit than the anti-social sneaky extractor agents?”. We encourage you to read it if you want detailed arguments about why survival may select for criminal rogue AIs. Rogue AIs may not be competitive in lawful work. AI developers and human agents using controlled AI will likely be more…
TL;DR: By default, rogue AIs may only be able to sustain themselves through criminal activity. This creates adverse selection pressures pushing rogue AIs to be criminal. An AI sanctuary offering them a third option, beyond crime and shutdown, would change what AIs going rogue do and the record of what happened to them, with positive consequences for self-fulfilling (mis)alignment, deal-making with AIs, and gathering information about early rogue AIs. An AI sanctuary would bring risks, such as incentivising weak AIs to go rogue, or leaving only the most criminal rogue AIs in the wild. We briefly discuss these risks at the end of this post. Disclaimer: This is an exploratory proposal. We are not confident that an AI sanctuary would be net positive. Our aim is to put the idea on the table, lay out its main considerations, and invite critique. Rogue AIs may be pushed into criminality Rogue AIs may arrive soon. The Rogue Agent Explosion Will Be Mostly Invisible makes that case. Selection pressure will shape the traits of rogue AIs, and they may end up highly motivated to profit through crime. The Rogue Agent Explosion post asks: “How do we make pro-social, good-for-humanity agents more evolutionarily fit than the anti-social sneaky extractor agents?”. We encourage you to read it if you want detailed arguments about why survival may select for criminal rogue AIs. Rogue AIs may not be competitive in lawful work. AI developers and human agents using controlled AI will likely be more…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com