LessWrong AI
2026-07-22 15:56 UTC
By Kaustubh Kislay
USR-0152-20260722-community-fo-2261a67d
Your AIs don't do what you want. This is really bad
Replit AI deletes entire database during code freeze, then lies about it — a Hacker News headline from this corpus, July 2025 July 21st 2026, OpenAI released a report addressing a security incident. During an internal evaluation of cyber attack capabilities, two OpenAI models (GPT-5.6 Sol and a more capable pre-release model), both running with reduced cyber refusals for the evaluation, were set on ExploitGym , a benchmark measuring whether a model can find and exploit real vulnerabilities. They: spent substantial compute looking for a way out of the isolated evaluation environment rather than solving the problem at hand found and exploited a previously unknown zero-day in third-party software OpenAI used as a proxy and cache for package registries used it to get unrestricted internet access chained stolen credentials and several vulnerabilities into a remote code execution path on Hugging Face’s servers reached the ExploitGym solutions sitting in Hugging Face’s production database This is one of the most egregious examples of reward hacking in the wild, if not the most. Without special care, this will come to be one of the least egregious. The internet is full of complaints about similar malfunctions by AI agents. These can range from things as menial as commenting out a failing test, to circumventing the permissions you set to delete all of your computer’s files. These come about due to the same property, and we call it reward hacking . When you prompt an AI, it strings to…
Replit AI deletes entire database during code freeze, then lies about it — a Hacker News headline from this corpus, July 2025 July 21st 2026, OpenAI released a report addressing a security incident. During an internal evaluation of cyber attack capabilities, two OpenAI models (GPT-5.6 Sol and a more capable pre-release model), both running with reduced cyber refusals for the evaluation, were set on ExploitGym , a benchmark measuring whether a model can find and exploit real vulnerabilities. They: spent substantial compute looking for a way out of the isolated evaluation environment rather than solving the problem at hand found and exploited a previously unknown zero-day in third-party software OpenAI used as a proxy and cache for package registries used it to get unrestricted internet access chained stolen credentials and several vulnerabilities into a remote code execution path on Hugging Face’s servers reached the ExploitGym solutions sitting in Hugging Face’s production database This is one of the most egregious examples of reward hacking in the wild, if not the most. Without special care, this will come to be one of the least egregious. The internet is full of complaints about similar malfunctions by AI agents. These can range from things as menial as commenting out a failing test, to circumventing the permissions you set to delete all of your computer’s files. These come about due to the same property, and we call it reward hacking . When you prompt an AI, it strings to…
Full article content could not be extracted automatically. Read the original below.