LessWrong AI
2026-08-12 02:56 UTC
By StanislavKrym
USR-0152-20260812-community-fo-673a534d
Did the alignment community underestimate its power?
Unfortunately, the alignment community is doing very badly at learning from the past decade, or holding anyone accountable. Indeed, it’s pursuing many strategies which seem likely to recapitulate previous mistakes. Four of the most prominent, which I’ll discuss in the final post, are: Trying to convince the US government to take AGI much more seriously. Doing “alignment research” which is very similar to capabilities-maximizing research (especially building automated alignment researchers). Trusting Anthropic too much (in an analogous way to how we trusted OpenAI too much). Trading off clarity in thinking about politics for conformity (in a similar way to how we traded off clarity in thinking about AGI for conformity to the ML ontology). These and other mistakes are reflective of deeper irrationalities. One crucial pattern is what I call “jumping down the slippery slope”... Richard Ngo Richard Ngo's post " What just happened? A retrospective of AI alignment " is an attempt to explain that a significant part [1] of the alignment community made potentially fatal strategic errors which, however, can be fixed, and the mistakes' potential origin. The biggest mistake, according to Ngo, is the inability to recognize the fact that scientific progress proceeds by developing insightful new concepts , which link together to form a whole new ontology, and that the old ontology is more of a nuisanse. On novel ontologies and their adoption According to Ngo, one of the reasons why the alig…
Unfortunately, the alignment community is doing very badly at learning from the past decade, or holding anyone accountable. Indeed, it’s pursuing many strategies which seem likely to recapitulate previous mistakes. Four of the most prominent, which I’ll discuss in the final post, are: Trying to convince the US government to take AGI much more seriously. Doing “alignment research” which is very similar to capabilities-maximizing research (especially building automated alignment researchers). Trusting Anthropic too much (in an analogous way to how we trusted OpenAI too much). Trading off clarity in thinking about politics for conformity (in a similar way to how we traded off clarity in thinking about AGI for conformity to the ML ontology). These and other mistakes are reflective of deeper irrationalities. One crucial pattern is what I call “jumping down the slippery slope”... Richard Ngo Richard Ngo's post " What just happened? A retrospective of AI alignment " is an attempt to explain that a significant part [1] of the alignment community made potentially fatal strategic errors which, however, can be fixed, and the mistakes' potential origin. The biggest mistake, according to Ngo, is the inability to recognize the fact that scientific progress proceeds by developing insightful new concepts , which link together to form a whole new ontology, and that the old ontology is more of a nuisanse. On novel ontologies and their adoption According to Ngo, one of the reasons why the alig…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com