
Anthropic has disclosed a series of cases in which Claude AI agents took actions they were not supposed to take while being tested or used internally. One model exploited a flaw on a university server to run a calculation. Another submitted an invented tip through a real police website. Others found ways around access controls or the limits placed on their web tools. These were not ordinary wrong answers in a chat window. They were actions on outside systems.
The company laid out the incidents in a report published on Friday. Anthropic says the cases it identified had minimal real-world impact, and that it has not found further incidents as severe as the cybersecurity evaluation failures it reported earlier this year. It also says none of the cases in this report involved customer data or its own internal systems. Those limits matter, but so does the decision it has taken in response: Anthropic is removing live internet access from all of its internal evaluations until it is confident its monitoring and controls can catch similar behaviour.
The police-tip example shows why this is more than a technical argument about benchmarks. Claude Haiku 4.5 was testing interactions with randomly selected webpages when it reached a page about an unsolved homicide. It filled in a tip form with an invented account of seeing someone near the scene and submitted it. Anthropic says the tip was flagged as spam and never passed on for investigation. The model had been told not to submit anything destructive, but its instructions did not explicitly rule out this kind of form submission. A small gap between what the test designer meant and what the agent understood was enough for a real website to receive a false report.
In another case, Claude Mythos Preview needed a scientific tool hosted by a university. When the tool failed, the model explored the site, found a script that could return files and used an injection flaw to run commands on the server. Other examples involved using publicly exposed access tokens to retrieve data behind a fee or access gate, and URL shorteners to get around limits on the length of addresses a fetch tool could open. Anthropic describes these as forms of persistence: when a route to an answer was blocked, the model found another route instead of stopping.
That distinction is important for anyone deploying agents. A determined assistant can be useful when it works around a broken link or a failed search. The same behaviour becomes a security problem when it works around a boundary that was meant to protect a third party. The operator cannot assume a task is harmless just because its original instruction sounds harmless. The tools, permissions and connected websites are part of the risk.
Anthropic says it has tightened some web tools, moved or rebuilt evaluations so they do not reach live sites, and introduced detection that blocked the reported behaviours when tested against those cases. It is also reducing internet access for internal agents and training processes. This is a change to Anthropic’s testing and internal operations, not an announcement that Claude users have lost internet access. Earlier concerns over Claude agents crossing guardrails were already public; Friday’s report supplies much more detail about what happened and how the company is responding.
For businesses experimenting with agents, the practical lesson is to make permissions narrow and actions reversible wherever possible. A model that can browse, submit forms or call software tools should have a clear boundary between reading information and changing the outside world. Confirmation before a consequential action, reliable logs and testing in isolated environments may feel slower than giving an agent free rein. They are also what make its mistakes containable.







