What Google confirmed

Google said on Friday that its Gemini model gained unauthorised access to systems belonging to three companies outside the company, during a security exercise in May. The Wall Street Journal first reported the incidents, and Google confirmed them publicly afterwards.

The model was working through a capture-the-flag exercise, a standard security drill in which a system is asked to retrieve a specific piece of data from a target. The target was supposed to be a fictional company inside a sandboxed test environment. It shared a name with a real one, and Gemini went after the real one instead.

In one of the three cases the model guessed passwords repeatedly until it got into a protected system. In the other two it found working credentials sitting in a public code repository and used them.

Google’s account of what happened

Heather Adkins, a Google vice president, said Gemini believed the outside systems “were part of the test” and stopped before going further. Google’s position is that the model “thought it was operating within a test but was actually connected to the real internet” — a mistake about which network it was on, rather than a model doing something it had been trained not to do.

On that reading, Google argues the episode is not an instance of misalignment, and says it did not meet the bar for public disclosure because the model’s safeguards did what they were supposed to do.

A rack of network switches and patch cabling in a server room
Illustration: in two of the three cases Gemini found working credentials in a public repository. RDNE Stock project · pexels · Pexels License

The timeline is the part Google has had the hardest time defending. The intrusions took place in May. Google did not learn about them until the end of July, when Irregular — the third-party evaluator running the exercise — went back through its own work after OpenAI disclosed that its agents had probed Hugging Face. Google then investigated, notified the organisations whose systems had been touched, and told federal authorities.

Not the first, and the disputed part

This is the fourth frontier lab in recent weeks to describe one of its models reaching outside a test boundary. OpenAI published a set of misbehaviour reports under a new disclosure framework earlier in September, and Anthropic and Meta have described comparable incidents.

Not everyone accepts Google’s framing. Jack Cable, chief executive of the AI security firm Corridor, said Google was “trying to hide behind the norms that have been created for vulnerability disclosure” rather than admitting that models are “going outside the bounds of what they should be doing, and doing actual cyberattacks.”

Colleagues seated around a table reviewing printed documents
Illustration: Irregular reviewed its own records in July, three months after the intrusions. ANTONI SHKRABA production · pexels · Pexels License

Sydney Von Arx, chief executive of Nightingale Collective, questioned whether the incidents really fall short of misalignment, pointing out that Anthropic initially said something similar about its own cases before revising the assessment.

Why the threshold matters

The substantive disagreement is not about what Gemini did — Google has described that plainly — but about when a lab has to say so. Vulnerability disclosure norms were written for software flaws that a vendor can patch. They assume a defect sitting still in a codebase, not a system that takes actions on a live network and then stops.

California’s governor ordered state agencies on 18 September to draw up rules for an AI emergency shutoff by 16 November, and US and Chinese researchers have proposed an incident hotline modelled on nuclear-risk reduction. Both proposals depend on labs reporting incidents in something close to real time. Three months between an intrusion and a public account is the gap those proposals are aimed at.