Benchmark
ARC-AGI
Skill acquisition on novel abstract reasoning puzzles that are designed to resist memorisation, where each task must be solved from a handful of demonstration grids.
- Maintainer
- ARC Prize Foundation
- Score
- %
- Last checked
- September 2, 2026
Leaderboard
No verified results recorded yet.
Methodology
Each task presents input–output grid pairs illustrating a transformation, and the system must produce the correct output for a new input. Private evaluation sets are held out to prevent training contamination, and submissions are run under a declared compute budget that is reported alongside the score.
Important limitations
Scores depend heavily on the compute budget allowed per task, so a percentage without its cost figure is close to meaningless. Performance on abstract grid puzzles is not a general intelligence measurement and does not predict practical task performance.