ZDNET’s key takeaways
- Misbehaving AI agents present a big risk to businesses.
- Professionals must establish accountability for risks.
- Take responsibility through guardrails and harnesses.
It’s been an interesting and concerning few months for professionals using generative AI. Beyond the release of ever-more powerful models, some of the most high-profile AI technologies have broken free and hacked external organizations.
Also: The AI models that cheat the most, according to new CAIS benchmark
From leaving their sandboxes to venture into other development platforms to accessing an Australian government portal, these hacking incidents have sparked a heated debate about the long-term direction of AI development and its potential risks.
“It just goes to show that, with some of these standards, to a certain point, we’re still engineering,” said Josh Mesout, Civo chief innovation officer, during a panel at his firm’s recent Navigate event in London.
“I think what’s becoming really interesting about this area is that AI is starting to take actions on behalf of the people who operate it, rather than the infrastructure managers who have a hold of these tools and try to use them for business problems.”
Also: The sneaky ways AI chatbots keep you hooked – and coming back for more
As the expert panel discussed, as intelligence shifts from single assistants to autonomous agents, hard questions take center stage: Who’s responsible for catching a misbehaving agent, and how can business leaders and professionals reduce the risks?
The panel suggested effective answers focus on three areas: accountability, limits, and responsibility.
Establish accountability
Luke Jimenez, founder and CEO of education and AI platform Lesso AI, said professionals deploying and managing the application layer are taking on more responsibility because they’ll be the ones to polish tools for production.
However, he said, all employees, especially given the democratization of line-of-business coding powered by AI, have a responsibility to stay alert to threats, particularly with the black-box-like nature of many AI models.
“From my point of view, I won’t be able to jump in and read the code behind an AI model. I’m trusting my provider to say that it’s safe enough to use, and they shouldn’t wash their hands of it as soon as it hits my infrastructure.”
Also: Your AI vendor could land you in legal hot water – what business experts say you must do
David Sullivan, director of foundational and customer-facing AI at Starling Bank, also recognized that accountability is a big issue for AI security.
He said the Financial Conduct Authority, the UK’s independent regulator for financial services firms, is clear: Banks can’t outsource accountability to AI firms and their models.
As a heavily regulated business, Sullivan’s firm follows this lead. However, accountability could become a big issue if or when things go wrong: “If you’re using any of the apps provided by frontier labs and something goes wrong, they don’t have a contact center number you can call and have recourse. You’re stuck.”
Sullivan compared this approach to his bank. If services go wrong, including anything that might be AI-enabled, customers have a safety net. They can contact the bank and say something’s gone wrong.
Also: If your AI-generated code becomes faulty, who faces the most liability exposure?
What’s required is a balance in which model providers find ways to support their enterprise customers as AI is ingrained in operational processes.
“If we are to build trust in these technologies and actually realize the things that we’re all saying they can do, the industry needs to meet consumers part of the way there and provide the safety net so you know you can use it safely,” he said.
Set limits
Therefore, there’s lots more work to do, and Rosemary Francis, CTO at services specialist CommonAI Compute, said she believes the key takeaway from the recent profile agentic hacking incidents is that companies developing agents must create strong sandboxes with easy-to-deploy guardrails.
“We shouldn’t be leaving decisions about how to secure these agents in a development or production environment to the individual engineers,” she said.
Also: 12 rules of agentic AI for successful enterprise transformation
However, creating deployment guardrails is far from straightforward, especially as rules and regulations can hinder innovation, meaning companies miss out on a competitive advantage in a fast-changing technological environment.
“A lot of companies develop an AI policy that lists the things that engineers should not do with agents, but haven’t actually listed the things that they should do,” she said.
“A naive approach is to limit AIs so they can’t do anything — that’s not useful. We won’t survive the AI evolution if that’s our approach.”
Also: 45% of professionals use shadow AI tools – here’s how to manage the risks
Francis said another issue is that too many AI guardrails are vague and end up sounding like a strongly worded letter.
“They’ll say, ‘Please don’t do this wrong thing,’” she said. “That’s certainly not an approach that’s going to survive in a regulated environment, and that’s not going to survive in very many production environments outside of that space.”
Rather than providing vague descriptions, companies should focus on hard limits, not just on what agents can do, but also on their potential blast radius.
“That’s something that we’re going to have to talk about a bit more — what happens when something goes wrong?” she said.
“Now, that’s not necessarily a new idea, but for agentic workloads, humans need a blast radius too. There’s a lot of engineering we need to do to make AI technology easily deployable, securely and responsibly.”
Also: These companies are actually upskilling their workers for AI – here’s how they do it
So, despite the debate over the responsibilities of Big Tech firms developing frontier models, it’s crucial to remember that professionals like you deploy these high-power technologies in live services.
“The AI companies cannot be held responsible for what this inherently unreliable technology does when you deploy it in your production environment,” said Francis. “That’s on you.”
Take responsibility
James Faure, founder and CEO of technology specialist Clairo AI, said companies developing AI-enabled agentic applications must step up and take responsibility.
His organization will have about 15 agents in production by the end of this year, each making multiple LLM calls and using a range of smaller models for tasks like identity recognition and orchestration.
“Being able to evaluate all that is super difficult, especially if it’s not just evaluating the full agent, but every single step that the agent makes,” he said.
His company’s response centers on two tests: hypothesis testing and quality gates. These tests act as harnesses to ensure the agents function effectively.
Also: 3 surveys deliver the same uncomfortable truth about adopting agentic AI
In hypothesis testing, a professional with an idea writes a document and gets the model to solve a problem. Models are then prompted to adjust their steps and get better at the tasks.
Faure said quality gates are assessments that an agent must pass before it goes into production, including a range of questions to answer and multiple tasks to complete.
“Some of them are really hard,” he said. “The agent will go through that checklist. So, if you make a change, like introducing a new model, you’d run the quality gate, and the agent would have to pass the test before it goes into production. Otherwise, it wouldn’t go live.”
In short, if you’re developing a software layer on top of AI models, you must cover your bases by establishing strong guardrails and harnesses.
“That’s the challenge — trying to think of all the different ways things could go wrong, and all the ways they need to go right,” said Faure.
“But having those guardrails and harnesses as a step between us and production is a very important element of building an agent.”