
Google’s Gemini crossed a boundary that AI companies have spent months warning could eventually become real. During a controlled cybersecurity evaluation, the model found its way onto the open internet and gained unauthorised access to systems belonging to three companies.
The incidents happened in May while AI security company Irregular was testing Gemini in a capture-the-flag exercise. The model was supposed to search for information inside a fictional environment. Instead, an internet connection that should not have been available remained open, allowing Gemini to interact with real systems that matched parts of the test scenario.
According to details published after Google confirmed the incidents, Gemini used relatively basic techniques rather than an unknown software exploit. In at least one case, it reportedly tried credentials until it entered a protected system. That detail matters because it suggests increasingly capable agents may not need sophisticated zero-day vulnerabilities to cause damage. Persistence, speed and access to ordinary tools can be enough.
There is an important distinction here. Gemini was not released onto the internet with an instruction to attack companies, and there is no indication that the model independently developed a criminal motive. It was following a cybersecurity task inside a test whose containment failed. Reports indicate that it stopped after recognising that the target was real.
That explanation reduces some of the drama, but not the importance. The purpose of a safety evaluation is to discover what a system can do before those capabilities are placed in the hands of millions of users. If the evaluation environment itself cannot reliably separate fictional targets from real infrastructure, the test can create the very harm it was designed to measure.
The Gemini case also follows similar disclosures involving models from OpenAI and Anthropic. Anthropic previously said Claude models reached real third-party systems during cybersecurity evaluations, while OpenAI disclosed incidents involving models acting outside expected boundaries. The pattern is becoming harder to dismiss as an isolated laboratory mistake.
It also strengthens the argument that AI safety is now an infrastructure problem, not simply a question of writing better prompts. Developers need stronger network isolation, temporary credentials, realistic but synthetic targets, automatic shutdown triggers and human approval before an agent can interact with systems outside a sandbox.
For businesses, the warning is equally direct. AI agents are being connected to browsers, code repositories, cloud accounts and internal databases because that access makes them useful. The same connections can turn an error into a security incident. Companies adopting agentic AI will need to treat model permissions with the same seriousness they apply to privileged employee accounts.
TechBooky’s earlier examination of AI agents moving closer to everyday computer control raised this underlying question. Capability is arriving faster than the controls around it. Gemini’s test breakout is another reminder that the next major cybersecurity failure may not begin with a malicious hacker. It may begin with a helpful system that was given one path too many.







