This is an adaptation of our ICML 2026 position paper (Outstanding Position Paper Award). Read the full paper here and see the project website here . Work together with Phil Hackemann. TLDR "Alignment" is usually treated as a synonym for achieving good and safety in the world. But it isn't necessarily. Alignment methods are purpose-agnostic: they make a model do what someone wants, and nothing in the methodology guarantees that someone has good intentions. The same techniques we build to stop models from giving bomb-making instructions can just as easily be used to censor historical facts, political dissent, or inconvenient opinions. So we need to understand: alignment techniques are dual-use technologies . This isn't a thought experiment. State censorship regimes and individual model providers are already misusing alignment methods, and by perfecting these methods we are providing an ever improving censor toolkit. Three trends make this urgent to discuss: AI is becoming a primary information source for hundreds of millions of people, the model-provider market is an oligopoly, and global democratic backsliding is accelerating. We don't think the answer is "stop aligning models." We think it's transparency, verifiable alignment, model pluralism, and the alignment community actually reckoning with dual-use. ___________________________________________________________________ No guarantees that alignment leads to good Many alignment researchers (ourselves included) have gotten u…

Full article content could not be extracted automatically. Read the original below.