DrivingBench is an evaluation with an unusual premise: it gives a frontier language model control of a real Toyota Corolla’s steering, accelerator and brakes, and asks it to drive a fixed cone course in an empty car park. The car moves at walking pace. A human sits in the driver’s seat ready to brake. The model sees the road through two cameras and drives using three tools — observe, set_motion and stop_now — issuing one command at a time.

On the current leaderboard, one model has finished. GPT-6 Astra reached 100 per cent progress with a finish time of 5:22, across three attempts that ran 49 per cent, then 100 per cent, then a did-not-finish. The successful attempt consumed 246.6 million tokens, which the benchmark costs at $7.74.

A view of an empty road from a camera mounted on a car dashboard
Models see the course through two cameras and issue one command at a time. Illustrative photograph. Sami Aksu · pexels · Pexels License

No other model has completed the course. Claude Fable 5.1 sits second with a best progress of 45 per cent, improving across its three attempts from 9 to 10 to 45 per cent. Grok 4.6 is third at 11 per cent, and GPT-5.6 Sol fourth at 6 per cent. All of those runs ended without finishing.

Why the gap is more interesting than the winner

A benchmark where the leader scores 100 and the runner-up scores 45 is not measuring a gradient. It is measuring whether a capability is present at all. The task combines spatial reasoning from camera images, continuous control under latency, and the discipline to keep issuing small corrections for five minutes without losing the thread — and on this course, three of the four models tested never got far enough to be scored on the last of those.

Progress is defined as how far along the course centreline an attempt travelled while staying within four metres of it, as a share of the centreline length to the finish zone. Distance, accepted commands and token cost are recorded alongside, with every command, the model’s stated reason for it, the car’s telemetry and the video aligned frame by frame.

An empty parking lot with painted markings seen from above
Progress is measured along the course centreline to the finish zone. Illustrative photograph. Jan van der Wolf · pexels · Pexels License

What it does not show

This is a low-speed cone course in a car park, not driving. There is no traffic, no pedestrians, no weather and no speed. The comparison to a purpose-built driving system is not close, and DrivingBench does not claim otherwise; the point is what a general-purpose model does when handed a physical control loop it was never trained for.

The attempt counts are also small — three runs per model — and the results are reported by the benchmark’s own operators rather than audited externally. A single 100 per cent run bracketed by a 49 and a DNF is a demonstration, not a reliability figure.

What to watch

Whether the second-place gap closes as labs ship models with better visual grounding, and whether DrivingBench publishes enough of its traces for someone outside the project to re-score a run.