01Leaderboard

Updated Aug 24, 2026

Mean effective score across 12 tasks. Rows with a 3-pass campaign also show Mean ± SEM and Attempt Pass rates; models without 3-pass leave those columns blank.

#AgentModelMean
1Claude CodeFable 5Anthropic89.2 ± 5.6
2Cursor CLIGrok 4.5xAI84.3 ± 7.9
3Claude CodeClaude Opus 5Anthropic83.1 ± 7.5
4Qoder CLIQwen3.8 MaxAlibaba81.4 ± 8.6
5Claude CodeClaude Opus 4.8Anthropic80.2 ± 7.5
6Qoder CLIKimi K3Moonshot77.3 ± 8.4
7Claude CodeClaude Sonnet 5Anthropic77.2 ± 7.4
8Gemini CLIGemini 3.6 FlashGoogle75.5 ± 9.5
9Qoder CLIGLM 5.2Zhipu74.0 ± 8.2
10Cursor CLIComposer 2.5Cursor72.6 ± 9.3
11Gemini CLIGemini 3.5 FlashGoogle72.1 ± 11.0
12Gemini CLIGemini 3.1 Pro PreviewGoogle70.9 ± 8.0
13Qoder CLIKimi K2.7 CodeMoonshot68.9 ± 8.8
14OpenCodeDeepSeek V4 FlashDeepSeek60.6 ± 10.3
15OpenCodeMiMo V2.5Xiaomi53.0 ± 10.6
16Claude CodeClaude Sonnet 4.6Anthropic52.6 ± 10.1
17Gemini CLIGemini 3.1 Flash-LiteGoogle46.1 ± 11.2
18OpenCodeDeepSeek V4 ProDeepSeek45.5 ± 11.4

* Auto is Cursor's built-in model router, not a fixed checkpoint. Pass@τ / Best-of-N are Attempt Pass rates over available 3-pass trials (not SWE-bench Pass@k). Single-run rows leave those columns as —.

01bCost / usage vs performance

Updated 2026-08-24

Each 3-pass model plotted by campaign cost or total token usage (log scale) against difficulty-weighted performance. The dashed line is the Pareto frontier — models no cheaper/leaner model beats.

vs

Score = difficulty-weighted mean across 12 tasks; Pass@0.5 = weighted Attempt Pass@0.5 over available (task, pass) attempts. Cost is the estimated 3-pass campaign spend; token usage is total input (incl. cache) over the 3-pass campaign.

01c3-pass charts

Mean reward ± SEM across 12 tasks, and Attempt Pass@1 (R ≥ 1) over available (task, pass) attempts for 3-pass campaign models only.

Mean reward R [%]

Attempt Pass@1 (R ≥ 1) [%]