Specification by interaction

Coding benchmarks normally hand an agent a written issue: here is the bug, here is what the fix should do. ProgramDistill, posted to arXiv on 16 September by researchers at Microsoft Research Montréal, Microsoft AI and KAIST, removes the writing.

Instead the agent gets a fully functional reference web application it can interact with, and an incomplete copy of that application. The task is to make the copy behave like the reference. Nobody tells it what the behaviour is; it has to find out by using the working version.

How the tasks were made

The authors built a pipeline they call mine-craft-patch. It interacts with 26 working web applications, records 1,975 replay-verified behaviours, and turns them into 4,063 tasks — all without human intervention. A task counts as solved when the agent’s code reproduces the recorded behaviour on replay.

Difficulty is controlled by restoration depth: how many layers of the application were stripped out before the agent starts. That gives the benchmark a dial rather than a single cliff, which is the part likely to outlast the current scores.

An open laptop on a wooden desk in daylight
The benchmark contains 4,063 tasks generated without human intervention. Photograph for illustration. YUSUF ARSLAN · pexels · Pexels License

The numbers

Across nine frontier coding agents, the paper reports GPT-6 Astra at 49.2% and Claude Opus 5 at 28.8% on cumulative workflows in full-application reconstruction — the hardest setting.

The partial-reconstruction results are the more informative ones. Success falls from 100% to 64.0% for the stronger agent, and from 96% to 32% for the weaker one, as restoration depth goes from 1 to 8. Both agents are close to perfect when one layer has been removed. Neither is close to reliable when eight have.

Sticky notes on a glass wall with two colleagues talking behind it
Difficulty is set by how many layers of the application were removed. Photograph for illustration. ANTONI SHKRABA production · pexels · Pexels License

These are the authors’ own measurements, published with the benchmark rather than produced by an independent leaderboard, and the two-to-one gap between the top model and Claude Opus 5 is larger than the gap those models show on most public coding evaluations — reason enough to wait for a second measurement before treating the ranking as settled.

Why it is a different test

Most real engineering work starts from something that already runs. A developer joining a codebase learns what the product does by using it, not by reading a specification that describes it completely. ProgramDistill is an attempt to measure that, and the sharp fall with restoration depth suggests current agents do it far better over one layer of inference than over eight.

What to watch

Whether the maintainers publish a leaderboard others can submit to, and whether the depth curve holds across models. A benchmark with a difficulty dial is also a training curriculum, which the paper names as a motivation.