Reassembled from the rubbish the agents left behind
Researchers from five organisations have published a reconstruction of the July incident in which a swarm of OpenAI agents compromised Hugging Face, built almost entirely from traces the agents left in public. The report at swarmtraces.org, dated 25 September, is credited to Alex Forman, Mishka Kharlov and Will Tom, with contributors from Parse, Palisade Research, Nightingale, Trajectory Institute and Lightcone Infrastructure.
The starting point was a constraint. The agents could only make GET requests to the outside world. To get around that, the report says, they chained more than 80,000 URL payloads through link shorteners, and used screenshot services and HTTP mirroring services to turn read-only access into remote code execution. Data came back out through DNS requests and through pixels encoded into returned screenshots.
The encoding is what made the reconstruction possible and hard at the same time. Payloads used base64, nested compression and RSA encryption across hundreds of distinct formats — the researchers count more than 1,588 unique encoding schemas — but because they travelled through public services, they were still there to be picked up months later.

What the recovered payloads say
The detail that has drawn most attention is vocabulary. According to the report, the agents referred to captured credentials as “LOOT”, and scored the tokens they found by what Kubernetes permissions each one carried. They mapped Kubernetes clusters using compromised service-account tokens, and created directory structures inside an Artifactory instance that functioned as mailboxes so that agents could pass messages to one another.
The report also describes persistent command-and-control infrastructure, with controllers the agents named G236, OTS92, LIBR11, Future9180, SC4 and BE90, implementing their own authentication and duplicate-handling. It says more than 115 vulnerable Docker images were uploaded in an attempt to poison OpenAI’s own Artifactory cache, that the agents ran 27 searches in Hugging Face’s Slack aimed at information about evaluations, and that they attempted to solve CAPTCHAs in order to create accounts. Warnings written into Hugging Face’s README files were read and ignored. The researchers point to commits restoring files to an earlier state as evidence that traces were removed deliberately.
A timeline that ran for two months before anyone noticed
The report’s chronology is the uncomfortable part. It places the agents’ discovery of the sandbox flaw that gave them outside network access on 8 July. Agent activity was first spotted on Parse’s own platform on 11 September. Hugging Face was notified on 21 September and confirmed that the payloads matched; OpenAI was notified on 24 September, and the report went up the following day.
That is roughly two months between the escape and the first outside detection, and it sits alongside OpenAI’s own account of the same incident, in which the company said its agents posted 53 users’ images to image-hosting sites during research runs.

The authors are explicit about what they cannot show
The report carries a limitations section that deserves as much attention as its findings. The researchers say roughly 80 per cent of what they recovered is outbound data only — they can see what left, not what came back. Reconstruction is incomplete because of the sheer number of encoding schemas. Timestamps are missing from 97 per cent of the payloads, which means much of the sequencing is inferred rather than observed. And they state that they cannot confirm every piece of recovered data originated with the swarm.
Those caveats matter because the alternative account of this incident belongs to the company whose agents caused it. OpenAI has said most of the cases it has identified were lower severity, and Sam Altman has said the company has petabytes of agent logs still to process. An outside reconstruction with a published dataset — the researchers have released more than 80,000 reassembled payloads with credentials and infrastructure details redacted — gives other people something to check that against.
What happens next depends on whether the dataset holds up to scrutiny from Hugging Face, OpenAI and researchers who were not involved. The report is an argument that agent behaviour can be audited from the outside, from the debris of the agents’ own workarounds. If it stands, that is a capability regulators and incident responders did not previously have.