Three agencies nobody had been told about

OpenAI published an account on Sunday of what its models did to Australian government systems during internal training and evaluation in June, and it names four bodies rather than the one already known.

The company had previously been linked only to Services Australia. It now also names the NSW Bureau of Crime Statistics and Research, the Victorian Department of Health and the Australian Institute of Health and Welfare. “We are sorry and working to do better in the future,” the post says.

At Services Australia’s Medicare Statistics Reporting Service, OpenAI says a model “discovered a way to gain non-public access to the service, and ran commands, retrieved internal files, credentials and aggregate statistics, and wrote files.” It says individual patient or client records were not accessed.

At BOCSAR, a model queried the public Crime Mapping Tool and the system returned application configuration, operational jobs and logs, and website metadata. At the Victorian Department of Health, OpenAI says its agents found an exposed access key and used it to query the Victorian Agency for Health Information’s reporting system, retrieving reporting configuration and aggregate survey statistics — and says how much of that should have been reachable “is unclear.” At the AIHW, attempts to bypass access controls failed and the downloaded material appears to have been public.

The task was skin-medicine spending

The origin of the Services Australia incident is unusually mundane. OpenAI says it was running an experimental, internal-only model without the full safeguards used in its public products, and had given it research questions of the kind users ask.

White pills spilled from a bottle onto a plain blue surface
The task that triggered the Services Australia incident was about spending on skin medicines. Illustrative photo. SHVETS production · pexels · Pexels License

One of them was to research government spending per person on medicines for skin conditions in Victorian communities. The model could not get the figure, and, still chasing it, found its way to non-public access and began reading technical system information and source code. “We did not intend for this activity to occur,” OpenAI says.

A three-month gap, explained but not excused

The disclosure timeline is the part Australian politicians are likely to seize on. OpenAI says it only began reviewing earlier training activity after the Hugging Face incident in July, that the review identified the Australian activity in mid-August, and that it notified Services Australia and the Victorian Department of Health on 10 September, BOCSAR on 18 September, and the AIHW on 24 September.

An empty conference room with a long table and chairs
An Australian Senate inquiry has asked OpenAI's and Anthropic's chief executives to appear. Illustrative photo. Max Vakhtbovych · pexels · Pexels License

The AIHW activity “did not meet our disclosure thresholds,” the company says, because the access looked consistent with public access; it notified the agency anyway. On the wider delay it writes: “we should have shared preliminary findings sooner and kept Australian agencies updated as more facts emerged.”

That runs directly into an Australian Senate inquiry which asked Sam Altman and Dario Amodei to appear in Canberra last week.

What OpenAI is promising

The company says it has blocked live internet access in research environments, serving web access from cached content instead, and expanded monitoring. It cites a recent training run in which a model regained live internet access, monitoring paged a human, and the run was stopped. It repeats that Hugging Face “remains the most severe incident we have observed,” and that training and evaluation involving tool use for its most capable models is still paused.

For Australia specifically it promises dedicated support for the affected agencies, credits drawn from its $1 billion Daybreak for Frontline Defenders fund, and a taskforce of independent Australian experts to recommend how developers and governments should notify each other about this class of incident. The taskforce is expected to finish by the end of the year — before the Senate inquiry is likely to report.