Benchmark

ARC-AGI

Skill acquisition on novel abstract reasoning puzzles that are designed to resist memorisation, where each task must be solved from a handful of demonstration grids.

Maintainer
ARC Prize Foundation
Score
%
Last checked
September 2, 2026

Leaderboard

No verified results recorded yet.

Methodology

Each task presents input–output grid pairs illustrating a transformation, and the system must produce the correct output for a new input. Private evaluation sets are held out to prevent training contamination, and submissions are run under a declared compute budget that is reported alongside the score.

Important limitations

Scores depend heavily on the compute budget allowed per task, so a percentage without its cost figure is close to meaningless. Performance on abstract grid puzzles is not a general intelligence measurement and does not predict practical task performance.