TL;DR In light of recent incidents, securing common AI research use cases needs a small set of building blocks that work together: hardened no-network sandboxes, real-time control monitors, monitoring-lifecycle infrastructure, and automated validation of security properties. Pieces of this exist. Nobody owns hardening them, making them secure by default, making them work together, fitting them to how research orgs actually operate, and keeping them working as models, frameworks and use cases change. By default we will get ad-hoc solutions rather than something well thought-out and it matters. This is a call for someone to step up and drive the effort. I can help connect you with funding opportunities and relevant people. Introduction The recent incidents ( OpenAI - HuggingFace , Anthropic , AISI ), where an agent with lowered safeguards either escaped a sandbox or attempted an attack on a 3rd party system, demonstrate we are now in a new regime: the AI models we study should be considered capable threat actors. Even if labs put in safeguards for the expected use, researchers often need to put models in contexts that increase the risk of misaligned and harmful actions. This requires appropriate mitigations - the alternative is either risking real harm, or missing out on important research and evaluations. A key assumption is we need to build measures effective against really strong models and agent swarms - at least a well-resourced top cyber offensive expert. Assuming anythi…

Full article content could not be extracted automatically. Read the original below.