LessWrong AI
2026-09-23 22:45 UTC
By Ashe Vazquez Nuñez
USR-0152-20260923-community-fo-0182c196
"I am an AI Safety Researcher"
Written as part of the MATS 9.1 extension program, mentored by Richard Ngo. Additional thanks to Andrew Wu, Maria Kostylew, and Lennie Wells for helpful draft feedback and editing. This post reflects on the tortured distinction between "safety" and "capabilities" in AI research. Richard Ngo has written about why the alignment vs. capabilities ontology is conceptually fraught , and is currently arguing that key strategic decision-makers in and around "AI safety" have brought about the AI labs' stampede [1] towards Artificial Superintelligence (ASI). This post instead looks at the following problem: how does one conduct alignment research without contributing to capabilities? It proposes decisions an individual or a small research group can take to do good work in AI. At the end, I discuss possible objections: namely, that my proposals fail to 'maximise impact'. I lay out why this meme is poisonous and usually backfires, and conclude by rejecting it entirely. Two examples of failure My first claim is that 'safety' and 'research' are two concepts that are in routine tension with one another. I illustrate this through examples of work that did too much of one at the expense of the other. Example: (mechanistic) interpretability In limiting its scope to remain innocuous, AI safety 'research' is habitually incurious and incrementalist. The last few years of interpretability serve as a good example. Interpretability's modern history originates from some cracked researchers noticing…
Written as part of the MATS 9.1 extension program, mentored by Richard Ngo. Additional thanks to Andrew Wu, Maria Kostylew, and Lennie Wells for helpful draft feedback and editing. This post reflects on the tortured distinction between "safety" and "capabilities" in AI research. Richard Ngo has written about why the alignment vs. capabilities ontology is conceptually fraught , and is currently arguing that key strategic decision-makers in and around "AI safety" have brought about the AI labs' stampede [1] towards Artificial Superintelligence (ASI). This post instead looks at the following problem: how does one conduct alignment research without contributing to capabilities? It proposes decisions an individual or a small research group can take to do good work in AI. At the end, I discuss possible objections: namely, that my proposals fail to 'maximise impact'. I lay out why this meme is poisonous and usually backfires, and conclude by rejecting it entirely. Two examples of failure My first claim is that 'safety' and 'research' are two concepts that are in routine tension with one another. I illustrate this through examples of work that did too much of one at the expense of the other. Example: (mechanistic) interpretability In limiting its scope to remain innocuous, AI safety 'research' is habitually incurious and incrementalist. The last few years of interpretability serve as a good example. Interpretability's modern history originates from some cracked researchers noticing…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com