ChatGPT-6 Astra was released on Sept 3, 2026 [1] . It demonstrates remarkably high benchmark scores across maths, scientific research and other domains. While OpenAI claims "Astra is our most aligned model" by showing 100% in ExploitBench and 0% in the ExploitGym honeypot [2] , the perfect score warrants closer scrutiny to what these numbers actually mean. In this article, I'll show that Astra is not sufficiently safety audited to be released to the public. Comparison to Mythos A jump in LLM capability usually comes from a fresh pre-training run, and occasionally a new architecture - we have seen such a jump before with Anthropic's Mythos [3] . Although its architecture was never published, Mythos was pre-trained from scratch with more training data and compute [4] . The same applies to Astra: with the new "looped transformer" architecture [5] , it had to be trained from scratch. This results in massive score jumps in benchmarks such as Terminal Bench Science , ARC-AGI-3 and FrontierMaths , as well as capabilities like trading, graphics rendering and coding. OpenAI and Anthropic, however, treat their safety measures very differently. Anthropic reached out to external auditors, including METR, UK AISI and other labs to red-team and test Mythos. They were given sufficient API access and time to conduct safety experiments. For example, METR spent three weeks red-teaming Anthropic's monitoring pipeline for Mythos. Furthermore, Anthropic's internal Frontier Compliance Framework e…

Full article content could not be extracted automatically. Read the original below.