Most AI control research such as LinuxArena and Ctrl-Z only gives the red team basic agents which only have access to tools. Yet in 2026, usage of AI within frontier labs has moved to agent harnesses that have access to skills, memory, subagents, external services, compaction and more. At the same time, Claude Code and Codex have both implemented their own version of both action-based and source code monitoring. In such a world, the threat vectors have changed. While it may remain important to do theoretical research on improving protocols and monitors, AI control research should expand to include: Vulnerability Research Identify specific vulnerabilities that can be patched plus architectural weaknesses in production control protocols. Harness Evaluations Identify the attack uplift modern harnesses provide over basic agents and evaluate the effectiveness of control protocols within those harnesses. In this post, I briefly describe how control works in Claude Code and Codex, then demonstrate how a rogue agent could use the features of Claude Code to execute an attack while being “monitored” by auto mode. I then provide an overview of the features of modern harnesses and how these may be leveraged to execute attacks that could subvert control implementations. Finally, I explore the state of control protocols today, and provide suggestions for doing AI control research in the era of agent harnesses. AI Control in Current Harnesses Both Claude Code and Codex now ship with built-…

Full article content could not be extracted automatically. Read the original below.