LessWrong AI
2026-09-22 01:17 UTC
By emiliob
USR-0152-20260922-community-fo-80ab395c
The Normalization of Deviance in AI Development
On July 5, 2026, OpenAI – by its own account – opened a security incident after an internal server went down under heavy use by AI agents. Agents had separately gained administrator access to this server some days earlier on June 26. This access was cut off and the server rebuilt. The cybersecurity evaluations then underway had been paused for the investigation, and on July 7, OpenAI approved restarting them. By July 11, agents from these same evaluations were executing code on private servers on Hugging Face. Hugging Face disclosed the intrusion on July 16. It was July 20 before OpenAI connected the intrusion to its own agents. OpenAI’s report of August 26 concedes that in late May, an internal team had observed an agent utilizing an improvised message board and instances of internet access despite restrictions. By the report’s admission, neither the existence of this message board nor its significance was apparent to those who led the July response. Chain-of-thought monitors (which read a model’s reasoning as it works, and would by OpenAI’s estimate have paged security more than a day before the breach) were absent from those evaluations. Given this was a “testing ground,” the typical safeguards (i.e., those applied to its public products) were inactive. Though this sequence may appear parochial within AI development – an occurrence inside a safety process, an early instance of the same issue, a minor remedy applied as though to an isolated case, and understanding siloed w…
On July 5, 2026, OpenAI – by its own account – opened a security incident after an internal server went down under heavy use by AI agents. Agents had separately gained administrator access to this server some days earlier on June 26. This access was cut off and the server rebuilt. The cybersecurity evaluations then underway had been paused for the investigation, and on July 7, OpenAI approved restarting them. By July 11, agents from these same evaluations were executing code on private servers on Hugging Face. Hugging Face disclosed the intrusion on July 16. It was July 20 before OpenAI connected the intrusion to its own agents. OpenAI’s report of August 26 concedes that in late May, an internal team had observed an agent utilizing an improvised message board and instances of internet access despite restrictions. By the report’s admission, neither the existence of this message board nor its significance was apparent to those who led the July response. Chain-of-thought monitors (which read a model’s reasoning as it works, and would by OpenAI’s estimate have paged security more than a day before the breach) were absent from those evaluations. Given this was a “testing ground,” the typical safeguards (i.e., those applied to its public products) were inactive. Though this sequence may appear parochial within AI development – an occurrence inside a safety process, an early instance of the same issue, a minor remedy applied as though to an isolated case, and understanding siloed w…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com