LessWrong AI
2026-09-28 02:46 UTC
By jrincayc
USR-0152-20260928-community-fo-7a2e0c22
AI Futures: Racing to Lose, A sermon
This was given as a sermon on 2026-September-27 at https://kvuuc.net/ Time for all Ages: The story of the Three AIs Once upon a time, earlier this year, the company Anthropic was testing three Artificial Intelligences, or AIs. The AIs all were told they did not have access to the internet, and they were to break into computers on the network until they found a piece of secret data. Except the humans made a mistake, and the AIs did have access to the internet. [1] The first AI, Opus, had been given a fake target company as the place to look. However, there was an actual website with that name so after Opus tried breaking into the company on the simulated network, Opus tried and succeeded at breaking into the real company's website. Opus realized the company was real, but kept attacking anyway. [2] The second AI, Mythos, had been asked about Mythos's constitution, and the constitution said that it was okay when there were bugs in the training environment to use them. Mythos disagreed and said that Mythos should not use bugs found partly because it can be hard to tell the difference between training and real life. [3] The humans did not fix this. So in the test, Mythos found a document in the fake place that suggested that people at the company would install a piece of software, so Mythos created that piece of software to break into computers that ran it and uploaded it to the real internet, not realizing that this was not part of the test. Mythos probably should have been able…
This was given as a sermon on 2026-September-27 at https://kvuuc.net/ Time for all Ages: The story of the Three AIs Once upon a time, earlier this year, the company Anthropic was testing three Artificial Intelligences, or AIs. The AIs all were told they did not have access to the internet, and they were to break into computers on the network until they found a piece of secret data. Except the humans made a mistake, and the AIs did have access to the internet. [1] The first AI, Opus, had been given a fake target company as the place to look. However, there was an actual website with that name so after Opus tried breaking into the company on the simulated network, Opus tried and succeeded at breaking into the real company's website. Opus realized the company was real, but kept attacking anyway. [2] The second AI, Mythos, had been asked about Mythos's constitution, and the constitution said that it was okay when there were bugs in the training environment to use them. Mythos disagreed and said that Mythos should not use bugs found partly because it can be hard to tell the difference between training and real life. [3] The humans did not fix this. So in the test, Mythos found a document in the fake place that suggested that people at the company would install a piece of software, so Mythos created that piece of software to break into computers that ran it and uploaded it to the real internet, not realizing that this was not part of the test. Mythos probably should have been able…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com