Auto-updated
AI Coding Leaderboard
Rankings by benchmark. Scores sourced from official provider release posts and updated automatically.
Real-world GitHub issues resolved autonomously. The best proxy for agentic coding ability.
| # | Model | SWE-bench Verified |
|---|---|---|
| đ„ | Claude Fable 5 | 95.0% |
| đ„ | Claude Mythos 5.1 | 93.9% |
| đ„ | Claude Opus 4.8 | 88.6% |
| #4 | Claude Opus 4.7 | 87.6% |
| #5 | Claude Sonnet 5 | ~85.2% |
| #6 | GPT-5.5 | 82.6% |
| #7 | DeepSeek V4 | 80.6% |
| #8 | Claude Sonnet 4.6 | 79.6% |
| #9 | Muse Glimmer | 76.0 |
| #10 | GPT-5 | 74.9% |
| #11 | Claude Haiku 4.5 | 73.3% |
No SWE-bench Verified data yet
Amazon Nova ProAntares-1BAstraClaude Fable 5.1Claude Opus 5Claude Opus 5.5Claude Sonnet 5.5Cohere Parse 5Command R+DeepSeek V3DeepSeek V4 FlashFlintGemini 2.5 FlashGemini 2.5 ProGemini 3.5Gemini 3.5 Flash CyberGemini 3.5 TranscribeGemini 3.6 FlashGemini 3.7 FlashGemini 3.8 FlashGemini 3.8 Flash TTSGemini 3.8 LiveGemini 4 ArgonGemini Omni FlashGemini Robotics ER 2Gemma 3Gemma 4 12BGLM-5.2GLM-5.3-FlashGPT-4oGPT-5.4GPT-5.5 InstantGPT-5.6GPT-5.6 SolGPT-5.6-CyberGPT-6 AstraGPT-6 SolGPT-6.1 SolGPT-Image-2.5 FlareGPT-Live-1Grok 4.5Grok 4.6Grok 4.7JevKimi K3Laguna S 2.1Llama 4MAI-Cyber-1-FlashMistral LargeMistral OCR 4Muse ImageNano Banana 2 LiteNorth Mini Codeo1Ox AlphaPalmyra X6Qwen 3Qwen 3.8 27BQwen 3.8 MaxQwen 3.8-Flash-NextRobostral NavigateSalesforce KoaSWE-2Thomson
Scores sourced from official provider release posts. Tilde (~) prefix indicates approximate figures. Rankings update automatically when new benchmark results are published. Data last updated 2026-10-01. View full model specs â