Private code, real tasks
A new coding benchmark from Specific Labs, a startup that says it is backed by Y Combinator, tests AI agents on tasks taken from private production codebases it licensed from real companies, rather than on public repositories a model may have seen during training. On Real-SWE’s first leaderboard, Anthropic’s Fable 5.1 running in Claude Code resolved 38.8% of its attempts. OpenAI’s GPT-6 Astra in Codex CLI followed at 33.8%, and Google’s Gemini 3.8 Flash in Gemini CLI at 31.2%.
The figures are Specific Labs’ own measurements, published in September 2026. They are independent of the model makers, but they have not been replicated by anyone else.
How the benchmark works
The tasks are problems that engineers at the source companies actually solved. Their instructions average 1,742 characters, and the median task requires edits to 11 files, according to the benchmark page. The companies include a competitor to Luma and Partiful with more than 200,000 users and a consumer fintech platform that processes more than 100,000 bank statements.
Each model gets eight independent attempts per task inside the agent harness its maker ships, so Real-SWE scores a model and its harness together rather than the model in isolation. The exception is Zhipu’s GLM 5.3, which ran in Claude Code. The page lists Snagnik Das, Siddhant Paliwal and Janak Sunil as its authors.

The leaderboard
| Rank | Model and harness | Resolved | Cost per attempt |
|---|---|---|---|
| 1 | Fable 5.1 (Claude Code) | 38.8% | $6.96 |
| 2 | GPT-6 Astra (Codex CLI) | 33.8% | $4.67 |
| 3 | Gemini 3.8 Flash (Gemini CLI) | 31.2% | $2.50 |
| 4 | GLM 5.3 (Claude Code) | 28.8% | $5.12 |
| 5 | Grok 4.6 (Grok Build) | 23.8% | $3.44 |
| 5 | Muse Spark 1.3 (Muse Code) | 23.8% | $2.74 |
| 7 | Kimi K3 (Kimi Code) | 18.8% | $3.90 |
| 8 | GPT-5.6 Sol (Codex CLI) | 16.2% | $2.65 |
The most common failure mode across all models was a missed requirement, the benchmark page says.
Cost per fix tells a different story
Dividing the cost of an attempt by the share of attempts that succeed changes the order. An analysis by The D*AI*LY Brief puts Fable 5.1 at about $17.94 per resolved task, GPT-6 Astra at $13.82, Muse Spark 1.3 at $11.51 and Gemini 3.8 Flash at $8.01.
Fable 5.1 is the most likely to finish a task. Gemini 3.8 Flash is the cheapest way to get a task finished, at less than half of Fable’s cost per fix, while resolving about four in five as many tasks.

Read the numbers with care
The sample is small. The public results rest on ten tasks, each run eight times for each of the eight models, according to The D*AI*LY Brief’s reading of the page’s rollout counts. That means a single task, solved or failed across all eight attempts, can move a model’s score by up to ten percentage points. The gap between second and third place is smaller than that.
The codebases are private by design, so outsiders cannot inspect the tasks. Sunil said on Hacker News that the team would “open source some of our tasks and model trajectories” and that it manually vets the codebases and the companies that supply them, the analysis reported.
What to watch
Whether Specific Labs publishes the tasks and trajectories it has promised, and whether the order holds as the task set grows beyond ten. A larger set would also show whether cheaper models keep their cost-per-fix advantage on harder tasks.