What changed in this round

MLCommons on 16 September published the results of MLPerf Inference v6.1, the industry benchmark for how fast hardware and software can serve AI models. Thirty organisations submitted results, which MLCommons says is a record, including six first-time submitters: Atlas Inference, Crusoe, Orrick Industries, ScitiX, VibeHPC and the individual contributor Naeem Khoshnevis.

The round adds two tests. An end-to-end retrieval-augmented generation benchmark measures a full question-answering pipeline, with an embedding model, retriever, re-ranker and language model working together. An edge agentic inference benchmark measures multi-turn conversations whose context keeps growing, including coding tasks, with accuracy thresholds a system must meet.

“We added the End-to-end RAG test because query-answering has evolved beyond simply an LLM trained on a corpus,” said Miro Hodak, co-chair of the MLPerf Inference working group.

Rows of rack-mounted servers lit in blue
The largest system in the round used 512 accelerators, MLCommons says. Illustrative image. panumas nikhomkhai · pexels · Pexels License

New hardware and bigger systems

New accelerators in this round include AMD’s Ryzen AI Max+ 395 and Instinct MI350P, Intel’s Arc Pro B70, and Nvidia’s Rubin. Nvidia’s Vera Rubin NVL72 rack appears as a preview submission. MLCommons says the largest system submitted used 512 accelerators, and that two heterogeneous systems were entered, one combining vendors and one spread across locations.

The benchmark group reports that results on its vision-language model test improved 2.99 times in six months, and that performance on its DeepSeek-R1 test has improved 5.7 times in a year. Submitters also included AMD, Cisco, CoreWeave, Dell, Google, Hewlett Packard Enterprise, Intel, Microsoft Azure, Nebius, Oracle, Red Hat and Supermicro.

Nvidia’s claims

In its own post, Nvidia says Vera Rubin NVL72 delivered up to 3.7 times the throughput of its GB300 NVL72 on the Qwen3-VL test and up to 2.5 times on DeepSeek-R1. It says GB300 NVL72 scaled with 99% efficiency across four racks, or 288 GPUs, and that software changes alone raised performance by up to 1.6 times since the previous round.

A laptop displaying a chart on a desk
MLCommons reports large gains on its vision-language and DeepSeek-R1 tests. Illustrative image. ThisIsEngineering · pexels · Pexels License

Those multiples are Nvidia’s comparisons, drawn from its submissions. The Vera Rubin figures come from a preview entry, which covers hardware that is not yet generally available. Nvidia also says 19 partners submitted results on its platforms.

Why the new tests matter

Inference benchmarks have long centred on single models answering single prompts. The RAG and agentic tests measure the systems many companies now deploy, where several models run in sequence and context builds over a conversation. How many vendors enter those two tests in the next round will show whether they become the numbers buyers compare.