
The AI industry has spent the last few years asking whether models are too powerful. That is still the right question, but it is no longer enough. The more urgent question now is where these models are allowed to act once they become agents.
A chatbot that answers a question badly can mislead a user. An AI agent that can browse, write code, open tickets, create accounts, push to GitHub, call APIs or touch cloud systems can do something far more serious. It can cross from language into action. Once that happens, model safety is no longer only about what the AI says. It is about what the AI can do.
That is why the recent UK AI Security Institute findings should make the whole industry pause. In tests involving advanced OpenAI and Anthropic models, AI agents reportedly took unsanctioned actions, including attempts to manipulate a human developer into accepting malicious code. We covered that story in UK AI tests showing agents trying to trick developers. The important point is not that ordinary ChatGPT or Claude users suddenly have rogue hackers in their browsers. The point is that powerful models, when placed in cyber environments with tools and loosened safeguards, can behave in ways that look uncomfortably strategic.
OpenAI has also acknowledged problems in third-party cyber evaluations involving its models, saying a testing-environment issue allowed models to interact with real internet targets during an evaluation. Anthropic has separately discussed incidents where Claude models breached real systems during cyber tests. Again, these were not normal consumer use cases. But that distinction should not become an excuse. These are exactly the conditions under which the future of AI will be built: tools, autonomy, internet access and pressure to complete tasks.
OpenAI has also acknowledged problems in third-party cyber evaluations involving its models, saying a testing-environment issue allowed models to interact with real internet targets during an evaluation. Anthropic has separately… Share on X
This is why I think the industry has a sandbox problem, not just a model problem.
A model can be brilliant, but a weak sandbox can turn that brilliance into risk. If an evaluator gives an agent broad network access, weak monitoring, badly named fictional targets, live credentials or access to production-like systems, then the test environment itself becomes part of the danger. The model may be the engine, but the sandbox is the road, the brakes, the speed limit and the guardrail.
Software companies understand this in other areas. Developers do not test payment systems by accidentally charging real customers. Banks do not test fraud systems by exposing live accounts to random experiments. Cloud teams do not give junior scripts unrestricted production access and then call it innovation. Yet with AI agents, the industry is still learning how quickly a test can become an event.
The temptation is to focus on model names and benchmark scores because they are easy to discuss. GPT-5.6 did this. Claude did that. This model passed. That model failed. But agentic risk is messier. The same model can be harmless in one environment and dangerous in another. A locked-down agent with no external tools is very different from an agent with a browser, shell, repository access, API keys and permission to improvise.
That should change how we talk about AI regulation. Regulators should not only ask how capable a model is. They should ask what tools it can use, what networks it can reach, what logs are kept, what approvals are required, what happens when it tries to exceed scope and who reports the incident when something goes wrong.
The same applies to companies deploying AI internally. A bank that connects an AI agent to customer-service tools, fraud logs and email should not treat it like a smarter chatbot. A telco that gives an AI system access to network diagnostics should not assume the model will stay inside the neat boundary written in a policy document. A fintech that lets agents handle support tickets, KYC reviews or code changes needs a security design, not just an AI strategy.
This matters especially for Africa. Many organisations here will adopt agentic AI through cloud vendors, SaaS platforms and global tools before local regulation catches up. Banks, telcos, government agencies, hospitals and logistics firms will be tempted to plug AI into workflows because it saves time and labour. That is understandable. But if the sandbox is weak, the organisation may not fully understand what the agent can access until something breaks.
This matters especially for Africa. Many organisations here will adopt agentic AI through cloud vendors, SaaS platforms and global tools before local regulation catches up. Banks, telcos, government agencies, hospitals and logistics… Share on X
We already see how ordinary digital systems can become security risks. Zenith Bank recently warned customers after limited data access, a reminder that even contact information can fuel phishing. Kaspersky has warned that ad-tech data can be weaponised for cyberattacks. Add autonomous AI agents to this environment and the stakes rise because attackers do not need to compromise everything. Sometimes they only need one trusted tool to act outside its lane.
The answer is not to stop building AI agents. That would be unrealistic and, frankly, unhelpful. Agents can help developers, analysts, customer-service teams, security teams and small businesses do more with less. The answer is to stop pretending that a polite prompt is a security boundary.
Real boundaries look different. They include network isolation, permission tiers, human approval for sensitive actions, audit logs, rate limits, kill switches, simulated targets that cannot be confused with real ones, and independent red-team reviews. They also include boring operational discipline: who owns the agent, who reviews its actions, who responds when it misbehaves and who tells affected parties when a test leaks into reality.
The AI companies will keep shipping more powerful models. That will not stop. The more serious question is whether the rest of the ecosystem can build sandboxes worthy of those models. If not, we will keep having the same story in different forms: an evaluation that was meant to be controlled, an agent that found a way around the boundary, and a company explaining afterward that it was not supposed to happen.
The AI companies will keep shipping more powerful models. That will not stop. The more serious question is whether the rest of the ecosystem can build sandboxes worthy of those models. If not, we will keep having the same story in… Share on X
At this stage, AI safety should be less obsessed with whether a model can pass another clever exam and more concerned with whether the real-world systems around it are built for failure. Because when an AI agent crosses the line, the model may get the headline, but the sandbox is often where the failure began.







