Anthropic ships Claude Fable 5.1 and cuts cache reads by 75%
The two models are the same system under different safeguard regimes. The pricing change, not the benchmark table, is what will show up in developers' bills.
Benchmark
The share of real GitHub issues from open-source Python repositories that a model resolves with a patch passing the project's own tests, on a human-validated subset of 500 tasks.
No verified results recorded yet.
A model receives an issue description and the repository, and must produce a patch. The patch is applied and the repository's existing test suite is run; a task counts as resolved only when the relevant tests pass. The Verified subset was human-reviewed to remove tasks that were unsolvable or ambiguously specified.
Coverage is limited to Python repositories present in the dataset, so results generalise imperfectly to other languages and codebases. Scaffolding, tool access and retry budget vary between submissions and materially change scores, so two numbers are only comparable when the harness is comparable.
Anthropic ships Claude Fable 5.1 and cuts cache reads by 75%
The two models are the same system under different safeguard regimes. The pricing change, not the benchmark table, is what will show up in developers' bills.
Privacy preferences
We use necessary cookies to run this site. With your consent we also use advertising cookies. You can change your choice at any time. Cookie Policy · Privacy Policy