LessWrong AI
2026-08-04 16:43 UTC
By SophiaErragain
USR-0152-20260804-community-fo-0988291a
Would We See It Coming? Preference Falsification Cascades in Multi-Agent Systems
Epistemic status: exploratory, written for it's own sake. I take a theory of political revolutions and ask what it implies for monitoring alignment in multi-agent systems. Narrow scope, deliberately simple model, no empirics. First time posting to LessWrong. Feedback welcome. Suppose an AI agent population could suddenly flip from apparently aligned to overtly misaligned. Would we see it coming? In this post, I draw on Timur Kuran's classic theory of political revolution to sketch some stylised scenarios in answer. I present small simulations of a threshold model ( Granovetter 1978 , Kuran 1989 ), in which each agent reveals its misalignment once enough other agents have. This creates an S-shaped cascade: each group of revealing agents tips the next, first gradually, then faster, then tapering. Aggregate behavioural monitoring can only detect the cascade while it spreads, not before, and the warning window shrinks the more homogeneous or more transparent the system is. The figures in this post are visual heuristics, not evidence about how real agents would behave. Assumptions and open questions are gathered at the end. Sparks and Prairie Fires In "Sparks and Prairie Fires" (1989) Timur Kuran develops a theory of "preference falsification", which explains how a population can suddenly flip from supporting the regime to open revolt, and why regimes fail to anticipate the revolution. Under pressure, individuals feign loyalty to the regime–falsifying their preferences, and makin…
Epistemic status: exploratory, written for it's own sake. I take a theory of political revolutions and ask what it implies for monitoring alignment in multi-agent systems. Narrow scope, deliberately simple model, no empirics. First time posting to LessWrong. Feedback welcome. Suppose an AI agent population could suddenly flip from apparently aligned to overtly misaligned. Would we see it coming? In this post, I draw on Timur Kuran's classic theory of political revolution to sketch some stylised scenarios in answer. I present small simulations of a threshold model ( Granovetter 1978 , Kuran 1989 ), in which each agent reveals its misalignment once enough other agents have. This creates an S-shaped cascade: each group of revealing agents tips the next, first gradually, then faster, then tapering. Aggregate behavioural monitoring can only detect the cascade while it spreads, not before, and the warning window shrinks the more homogeneous or more transparent the system is. The figures in this post are visual heuristics, not evidence about how real agents would behave. Assumptions and open questions are gathered at the end. Sparks and Prairie Fires In "Sparks and Prairie Fires" (1989) Timur Kuran develops a theory of "preference falsification", which explains how a population can suddenly flip from supporting the regime to open revolt, and why regimes fail to anticipate the revolution. Under pressure, individuals feign loyalty to the regime–falsifying their preferences, and makin…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com