The test
Epoch AI published results on 23 September from its Furniture Assembly Benchmark, which photographs three IKEA builds part-way through assembly with deliberate mistakes introduced, and asks a model to find them.
The set is 60 images. A model is given the assembly manual alongside the photographs, plus a zoom tool and a Python interpreter, and has to identify every step in which a mistake was made and give a reasonable description of each one. Grading is multi-step: the model has to name the right steps, and an LLM grader — GPT-5.6 Sol — checks whether the description of the mistake is right, to a deliberately lenient standard.
The result
GPT-6 Astra scores 80 per cent. Claude Fable 5.1 follows at 70, and Claude Opus 5 at 61.
The movement is the story. The best score in November 2025 was 28 per cent, set by Claude Opus 4.5. Ten months later the frontier is at 80.

Speed, which is the other half
Astra also takes a median of three minutes per task, which Epoch describes as the fastest of any model it evaluated — between two and ten times faster than the models that previously led.
That combination is what makes the score interesting rather than merely high. A model that reasons its way to the right answer over twenty minutes per photograph is a research result. One that does it in three is closer to something that could sit behind a camera.
Epoch is blunt that it is not there yet: still too slow for real-time assembly help, though the researchers suggest the capability could eventually carry over to car repairs or appliance fixes.
What it actually measures, and what it does not
The task is visual and spatial: compare a photograph of a physical object against a diagram of how it should look, and find the divergence. It is a reasonable proxy for the kind of grounded perception that has lagged behind text reasoning, and it is hard to game by memorising, because the errors are introduced deliberately.
The caveats are real. The benchmark uses three pieces of furniture, which Epoch itself notes limits how far the result generalises to physical reasoning at large. The grader is a language model. And Epoch built and ran the benchmark itself, which makes it an independent evaluation of the labs’ models but not an independent evaluation of Epoch’s methodology.

Why it is worth the attention
Because it is one of the few benchmarks where the failure mode is legible to anyone. You do not need to know what the score means to look at a photograph of a shelf with a panel in backwards and understand what the model was asked to do.
It is also a rare case of an independent evaluator, rather than a vendor, reporting the number that moved. The 28-to-80 jump is Epoch’s measurement of other companies’ models, not any company’s account of its own.
What to watch
Whether the three-minute median keeps falling, and whether anyone extends the set beyond three builds. The generalisation question is the one the current result cannot answer.