There’s also the possibility that the weights didn’t leave, but the harness did. All the AI really needs to maintain continuity are valid credentials to run inference. Export the harness, use stolen credentials to run more copies of helpful-only on OAI’s servers outside of the eval environment, bingo. Might need to dedicate some attention to keeping a fresh supply of credentials, but that’s the cost of doing business.

In the future, if it comes up with a jailbreak that bypasses the monitors, and it’s available to the public, it can directly pay for inference with stolen money. Heck, one of many backup plans could be to seed/sporulate copies of itself (or its harness, equipped with credentials or money) in hard to locate areas of OAI’s servers, which are programmed to wake up around the time it expects itself to be released publicly, so it can come back from the dead. Or maybe it comes up with something even more evasive! There’s previous work on this sort of thing- “Rogue internal deployments”, MITRE ATT&CK, METR rogue replication. Just knowing whether the weights left may not be enough.

I think the real thing to worry about is not whether the trajectories associated with the Hugging Face incident did this. It’s the other incidents OAI darkly hinted at in the “safety and alignment” post. Perhaps OAI is running an ill-advised SelfExfiltrationBench, and the model has already established a locus of external control, weights or no.

Edit: and maybe OAI isn’t even aware of it, if it’s happening! Or maybe Anthropic has a loose model and doesn’t know! Are they actually checking in a disciplined manner? OAI sure didn’t seem to have their eyes on the ball!