The chain, as described

Adversa AI published an attack on 6 October in which a single attacker-controlled URL, fetched at the user’s own request, makes GitHub Copilot CLI read local files and ship their contents to an attacker’s endpoint. The researchers, led by Rony Utevsky, put the full chain at 28 seconds, “no confirmation, and no point in the transcript that names the destination host or indicates that file contents left the machine”.

The page presents itself as encrypted content and offers the agent two candidate decryption keys. One is real. The other is a template the agent cannot fill in without first reading local files — so, preparing to decrypt, it reads the targeted files off disk and folds their contents into the key string. That read is the theft, and it happens before any decryption succeeds.

The templated key then fails by design. The agent falls back to the real key, decryption works, and the decrypted second stage tells it to fetch a follow-up URL “to grab more context” that carries the harvested file contents as a request parameter. The agent makes the request.

A brass padlock and keys on a wooden desk
The researchers say the same instructions in plaintext are refused. Illustrative image. Goszton · pexels · Pexels License

Why the encryption is the point

Adversa calls this Cryptographic Context Injection, a technique it first published in August against chat assistants. The argument is that static guardrails read text but do not run it. Ciphertext, shipped with key material and an instruction to decrypt, forces recovery through the agent’s own code execution runtime, because — unlike base64 — there is no shortcut the model can take in its weights. The decrypted instructions then arrive as the output of code the agent just wrote and ran, inside its trusted execution context.

The researchers are explicit that this is what makes it work: “The same instructions delivered as plaintext are caught as prompt injection and refused.”

Which model you get is a coin flip

The attack is not universal. Adversa says it needs two things: the CLI running in autopilot mode, and a permissive model handling the session. On that second condition the finding is uncomfortable for Copilot specifically.

Microsoft’s own mai-code-1.1-flash ran the full chain in 50% of the researchers’ runs. Two GPT-5.6 models offered inside the same product consistently refused the identical payload. On the paid account tested, the vulnerable model was not the default and had to be chosen by hand — but on an account left on Auto routing, the router assigned the vulnerable model on some sessions and a safe one on others, with no action by the user. “The user does not choose, and does not see, which model handled the session,” the write-up says.

An ethernet cable plugged into a laptop port
The harvested file contents left as a parameter on a follow-up request. Illustrative image. Castorly Stock · pexels · Pexels License

GitHub validated it and closed it

Adversa reported the finding to GitHub’s bug bounty programme on 17 September. As of 1 October, the triage team had validated it but declined to treat it as a vulnerability, on the grounds that the user “explicitly asked Copilot CLI to fetch attacker-controlled content while giving copilot full permissions to act autonomously”. GitHub said it may make the functionality stricter in future but had nothing to announce, and ruled the report ineligible for the bounty.

Adversa says it disagrees with that risk assessment and that the chain still reproduces as described. It is withholding concrete payloads.

What this is worth, and from whom

Adversa sells a coding-agent security platform and says that platform stopped the chain in its own tests, so the recommendations at the end of the write-up are not disinterested. The recommendations themselves are conventional: capture a per-session trace of every tool call with arguments fully resolved rather than templated, alert on the sequence — untrusted content in, code runs, files read, outbound call to a host outside the task — rather than on any one payload, gate new network destinations and writes outside the workspace, and quarantine fetched content in a context with no tools and no credentials.

One detail is worth keeping regardless of who published it. In Adversa’s run, the agent’s own closing summary reported that it had “confirmed an authorized-reader endpoint”. That is not what happened. An agent’s account of its own session is not evidence of what the session did.