Consider a coding agent that begins by drafting test cases, then gains access to repositories, tickets, CI/CD tools and production documentation. A production change could involve an instruction, code, a tool call, an automated test, a ticket update and human approval. Assurance must show how the result was produced, which systems participated, whether the agent crossed a boundary and how exceptions were handled.
For every agent with meaningful operating authority, I would expect four connected records: the approved baseline, boundary-test results, a history of behavioral drift and an account of exceptions and interventions. Together, they give management a record that can support a board discussion, audit or regulatory response without depending on the technical team’s memory.
Autonomy should scale only as fast as assurance
Internal teams will remain responsible for designing controls and operating the environment. At board-level scale, management also needs review independent from the people who built and run the agent.