01Leaderboard

Updated Jun 28, 2026

Mean effective score across all 12 tasks. Source: Table 3 in the paper.

#AgentModelMean score
1Claude CodeClaude Opus 4.7Anthropic90.1%
2CodexGPT-5.5OpenAI86.1%
3Claude CodeClaude Opus 4.8Anthropic85.7%
4Cursor CLIAuto* · routerCursor81.1%
5Gemini CLIGemini 3.5 FlashGoogle78.1%
6Cursor CLIComposer 2.5Cursor74.6%
7OpenCodeDeepSeek V4 FlashDeepSeek65.5%
8OpenCodeMiMo V2.5 ProXiaomi60.9%
9OpenCodeMiMo V2.5Xiaomi60.6%
10Claude CodeClaude Sonnet 4.6Anthropic59.2%
11OpenCodeDeepSeek V4 ProDeepSeek49.8%
OpenCodeGLM 5.2Zhipuin eval
Kimi CodeKimi 2.7 CodeMoonshotin eval

* Auto is Cursor's built-in model router, not a fixed checkpoint. Kimi 2.7 Code and GLM 5.2 evaluations are in progress.