A preprint posted to arXiv on 2 September claims the first AI system to outscore the top-ranked human competitor on an International Olympiad in Informatics problem set.

The numbers

At IOI 2026 the authors report their competition-specific system scoring 535.4 points out of 600, against a gold threshold of 361.12 and a top human score of 498.27.

Run against the 2025 problem set, the smaller of two models — Nemotron-3-Nano-CC, at 30B parameters with 3B active — improved from 130 points to 468 once the authors applied a test-time strategy they call GenCorrect. The gold threshold that year was 438.3. The larger Nemotron-3-Ultra-CC, 550B with 55B active, scored 502.

A whiteboard covered in mathematical equations.
GenCorrect is applied at inference rather than during training, and accounts for most of the smaller model's gain. cottonbro studio · pexels · Pexels License

How it was built

The pipeline combines curated problems, synthetic reasoning traces, supervised fine-tuning and reinforcement learning, over a set of 22,000 curated problems. GenCorrect is applied at inference rather than during training, which is where most of the Nano model’s gain comes from: 130 to 468 is not a better model, it is a better way of using the same one.

The caveats the paper carries

This is a preprint. It has not been peer reviewed, and the results have not been independently reproduced.

The IOI 2026 figure comes from what the authors describe as a competition-specific system rather than a general model, which makes the comparison against a human competitor less direct than the headline suggests. A contestant sits an exam under fixed conditions; a system tuned for that exam is a different kind of entrant.

A person typing on a laptop at a desk late at night.
The work is a preprint: not peer reviewed, and not independently reproduced. Faye Tsui · pexels · Pexels License

What it does say

That a 30-billion-parameter model with 3 billion active can be pushed past a gold threshold by inference-time technique alone is the more durable result, and the one least dependent on how the human comparison is framed.

What to watch

Whether the IOI itself comments. Competitive programming has handled AI entrants informally so far, and a claim to have beaten the top human is the kind of thing that forces a rule.