01Leaderboard
Updated Aug 24, 2026Mean effective score across 12 tasks. Rows with a 3-pass campaign also show Mean ± SEM and Attempt Pass rates; models without 3-pass leave those columns blank.
| # | Agent | Model | Mean | ||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Code | Fable 5Anthropic | 89.2 ± 5.6 | 77.8% (28) | 91.7% (33) | 83.3% (10) | |
| 2 | Cursor CLI | Grok 4.5xAI | 84.3 ± 7.9 | 75.0% (27) | 86.1% (31) | 83.3% (10) | |
| 3 | Claude Code | Claude Opus 5Anthropic | 83.1 ± 7.5 | 72.2% (26) | 86.1% (31) | 75.0% (9) | |
| 4 | Qoder CLI | Qwen3.8 MaxAlibaba | 81.4 ± 8.6 | 69.4% (25) | 77.8% (28) | 75.0% (9) | |
| 5 | Claude Code | Claude Opus 4.8Anthropic | 80.2 ± 7.5 | 63.9% (23) | 86.1% (31) | 75.0% (9) | |
| 6 | Qoder CLI | Kimi K3Moonshot | 77.3 ± 8.4 | 66.7% (24) | 77.8% (28) | 75.0% (9) | |
| 7 | Claude Code | Claude Sonnet 5Anthropic | 77.2 ± 7.4 | 63.9% (23) | 80.6% (29) | 75.0% (9) | |
| 8 | Gemini CLI | Gemini 3.6 FlashGoogle | 75.5 ± 9.5 | 66.7% (24) | 72.2% (26) | 75.0% (9) | |
| 9 | Qoder CLI | GLM 5.2Zhipu | 74.0 ± 8.2 | 61.1% (22) | 77.8% (28) | 75.0% (9) | |
| 10 | Cursor CLI | Composer 2.5Cursor | 72.6 ± 9.3 | 58.3% (21) | 75.0% (27) | 66.7% (8) | |
| 11 | Gemini CLI | Gemini 3.5 FlashGoogle | 72.1 ± 11.0 | 63.9% (23) | 66.7% (24) | 66.7% (8) | |
| 12 | Gemini CLI | Gemini 3.1 Pro PreviewGoogle | 70.9 ± 8.0 | 52.8% (19) | 72.2% (26) | 58.3% (7) | |
| 13 | Qoder CLI | Kimi K2.7 CodeMoonshot | 68.9 ± 8.8 | 55.6% (20) | 72.2% (26) | 75.0% (9) | |
| 14 | OpenCode | DeepSeek V4 FlashDeepSeek | 60.6 ± 10.3 | 47.2% (17) | 63.9% (23) | 66.7% (8) | |
| 15 | OpenCode | MiMo V2.5Xiaomi | 53.0 ± 10.6 | 38.9% (14) | 61.1% (22) | 50.0% (6) | |
| 16 | Claude Code | Claude Sonnet 4.6Anthropic | 52.6 ± 10.1 | 36.1% (13) | 52.8% (19) | 50.0% (6) | |
| 17 | Gemini CLI | Gemini 3.1 Flash-LiteGoogle | 46.1 ± 11.2 | 33.3% (12) | 50.0% (18) | 41.7% (5) | |
| 18 | OpenCode | DeepSeek V4 ProDeepSeek | 45.5 ± 11.4 | 36.1% (13) | 52.8% (19) | 50.0% (6) |
* Auto is Cursor's built-in model router, not a fixed checkpoint. Pass@τ / Best-of-N are Attempt Pass rates over available 3-pass trials (not SWE-bench Pass@k). Single-run rows leave those columns as —.
01bCost / usage vs performance
Updated 2026-08-24Each 3-pass model plotted by campaign cost or total token usage (log scale) against difficulty-weighted performance. The dashed line is the Pareto frontier — models no cheaper/leaner model beats.
Score = difficulty-weighted mean across 12 tasks; Pass@0.5 = weighted Attempt Pass@0.5 over available (task, pass) attempts. Cost is the estimated 3-pass campaign spend; token usage is total input (incl. cache) over the 3-pass campaign.
01c3-pass charts
Mean reward ± SEM across 12 tasks, and Attempt Pass@1 (R ≥ 1) over available (task, pass) attempts for 3-pass campaign models only.