Auto-updated
AI Coding Leaderboard
Rankings by benchmark. Scores sourced from official provider release posts and updated automatically.
Real-world GitHub issues resolved autonomously. The best proxy for agentic coding ability.
| # | Model | SWE-bench Verified |
|---|---|---|
| š„ | Claude Fable 5 | 95.0% |
| š„ | Claude Opus 4.8 | 88.6% |
| š„ | Claude Opus 4.7 | 87.6% |
| #4 | Claude Sonnet 5 | ~85.2% |
| #5 | GPT-5.5 | 82.6% |
| #6 | DeepSeek V4 | 80.6% |
| #7 | Claude Sonnet 4.6 | 79.6% |
| #8 | Muse Glimmer | 76.0 |
| #9 | GPT-5 | 74.9% |
| #10 | Claude Haiku 4.5 | 73.3% |
No SWE-bench Verified data yet
Amazon Nova ProAntares-1BClaude Opus 5Command R+DeepSeek V3DeepSeek V4 FlashFlintGemini 2.5 FlashGemini 2.5 ProGemini 3.5Gemini 3.5 Flash CyberGemini 3.6 FlashGemini 3.7 FlashGemini Omni FlashGemini Robotics ER 2Gemma 3Gemma 4 12BGLM-5.2GPT-4oGPT-5.4GPT-5.5 InstantGPT-5.6GPT-5.6 SolGPT-5.6-CyberGPT-Live-1Grok 4.5Grok 4.6Kimi K3Laguna S 2.1Llama 4MAI-Cyber-1-FlashMistral LargeMistral OCR 4Muse ImageNano Banana 2 LiteNorth Mini Codeo1Palmyra X6Qwen 3Qwen 3.8 27BQwen 3.8 MaxRobostral Navigate
Scores sourced from official provider release posts. Tilde (~) prefix indicates approximate figures. Rankings update automatically when new benchmark results are published. Data last updated 2026-08-21. View full model specs ā