AI Models Compared
Every major frontier model - pricing, context window, benchmark scores, and what each one is actually best at. Continuously web-verified; last checked 2026-08-20.
Pricing at a glance
| Model | Provider | Strong at | Context | Input / 1M | Output / 1M | Type |
|---|---|---|---|---|---|---|
| DeepSeek V4 FlashPreview | DeepSeek | Long context | 1M tokens | Free | Free | Open |
| DeepSeek V4Preview | DeepSeek | Long context | 1M tokens | Free | Free | Open |
| GPT-5.5 | OpenAI | Coding | 1,050,000 tokens (128,000 max output) | $5.00 | $30.00 | Closed |
| Claude Opus 4.8 | Anthropic | Coding | 1M | $5.00 | $25.00 | Closed |
| Claude Sonnet 4.6 | Anthropic | Long context | 1M | $3.00 | $15.00 | Closed |
| Claude Opus 4.7 | Anthropic | Coding | 1M | $5.00 | $25.00 | Closed |
| Claude Haiku 4.5 | Anthropic | Speed | 200K | $1.00 | $5.00 | Closed |
| GPT-5 | OpenAI | Math | 400,000 tokens (128,000 max output) | $1.25 | $10.00 | Closed |
| GPT-4o | OpenAI | Speed | 128,000 tokens (16,384 max output) | $2.50 | $10.00 | Closed |
| o1 | OpenAI | Reasoning | 200,000 tokens (100,000 max output) | $15.00 | $60.00 | Closed |
| Gemini 3.5 | Google DeepMind | Coding | 1,048,576 tokens (Gemini 3.5 Flash; Pro variant not yet released) | $1.50 | $9.00 | Closed |
| Gemini 2.5 Pro | Google DeepMind | Long context | 1,048,576 tokens (1M) input; up to 65K output | $1.25 | $10.00 | Closed |
| Gemini 2.5 Flash | Google DeepMind | Cost efficiency | 1,048,576 tokens (1M) input; up to 65,535 output | $0.30 | $2.50 | Closed |
| Amazon Nova Pro | Amazon Web Services | Multimodal | 300K tokens | $0.80 | $3.20 | Closed |
| DeepSeek V3 | DeepSeek | Cost efficiency | 128K tokens | Free | Free | Open |
| Llama 4 | Meta | Long context | Up to 10M tokens (Scout); ~1M tokens (Maverick) | Free | Free | Open |
| Qwen 3 | Alibaba (Qwen Team) | Cost efficiency | 128K tokens (32K for 0.6B/1.7B/4B dense variants) | Free | Free | Open |
| Mistral Large | Mistral AI | Reasoning | 128000 | 2.00 | 6.00 | Closed |
| Gemma 3 | Google DeepMind | Cost efficiency | 128K tokens (32K for the 1B variant) | Free | Free | Open |
| Command R+ | Cohere | Long context | 128K tokens | $2.50 | $10.00 | Closed |
| Gemma 4 12B | Cost efficiency | 256K tokens | Free | Free | Open | |
| Claude Fable 5 | Anthropic | Coding | 1M | $10.00 | $50.00 | Closed |
| North Mini Code | Cohere | Cost efficiency | 256K tokens | 0 | 0 | Closed |
| GPT-5.4 | OpenAI | Long context | 1,050,000 tokens (128,000 max output) | $2.50 | $15.00 | Closed |
| GPT-5.5 Instant | OpenAI | Speed | 1050000 | 5.00 | 30.00 | Closed |
| Mistral OCR 4 | Mistral AI | Cost efficiency | Not announced | $4.00 per 1,000 pages (API) / $2.00 per 1,000 pages (Batch) | $5.00 per 1,000 pages (Document AI annotation tier) | Closed |
| GPT-5.6 | OpenAI | Coding | 1.05M | 5 | 30 | Closed |
| GPT-5.6 SolPreview | OpenAI | Coding | 1050K | $5.00 | $30.00 | Closed |
| Claude Sonnet 5 | Anthropic | Coding | 1M | $3.00 | $15.00 | Closed |
| FlintPreview | Springboards | Speed | Not announced | Not announced | Not announced | Closed |
| Nano Banana 2 Lite | Speed | 1M | 0.25 | 1.50 | Closed | |
| GLM-5.2 | Zhipu AI / Z.ai | Long context | 1M | $1.40 | $4.40 | Open |
| Muse Image | Meta | Cost efficiency | Not announced | Not announced | Not announced | Closed |
| GPT-Live-1 | OpenAI | Speed | Not announced | Not announced | Not announced | Closed |
| Grok 4.5 | SpaceXAI (xAI) | Cost efficiency | 500K | $2.00 | $6.00 | Closed |
| Robostral Navigate | Mistral AI | Cost efficiency | Not announced | Not announced | Not announced | Closed |
| Gemini Omni FlashPreview | Multimodal | 1000000 | 1.50 | 17.50 | Closed | |
| Kimi K3 | Moonshot AI | Coding | 1M | $3.00 | $15.00 | Open |
| Gemini 3.6 Flash | Speed | 1M | $1.50 | $7.50 | Closed | |
| Laguna S 2.1 | Poolside | Coding | 1M | Not announced | Not announced | Open |
| Qwen 3.8 Max | Alibaba | Long context | 983616 | 2.00 | 6.00 | Closed |
| Gemini 3.5 Flash CyberPreview | Coding | 1000000 | 1.50 | 7.50 | Closed | |
| Antares-1B | Cisco | Cost efficiency | 128000 | Free | Free | Open |
| Claude Opus 5 | Anthropic | Long context | 1M | $5.00 | $25.00 | Closed |
| MAI-Cyber-1-FlashPreview | Microsoft | Coding | 256K | Not announced | Not announced | Closed |
| Gemini Robotics ER 2 | Google DeepMind | Multimodal | 128K | $2 | $10 | Closed |
| GPT-5.6-CyberPreview | OpenAI | Coding | 1.05M | $12.50 | $75.00 | Closed |
| Muse Glimmer | Meta | Cost efficiency | 128K | Free | Free | Open |
| Palmyra X6 | Writer | Cost efficiency | 1000000 | $2.00 | $8.00 | Closed |
| Gemini 3.7 Flash | Long context | 1M | $0.75 | $3.75 | Closed | |
| Grok 4.6 | SpaceXAI | Cost efficiency | 500K | $2.00 | $6.00 | Closed |
| Qwen 3.8 27B | Alibaba | Coding | 262K | Free | Free | Open |
Prices in USD. Open-source models are free to self-host; API pricing varies by provider.
Closed-source models
GPT-5.5
1,050,000 tokens (128,000 max output) ctxOpenAI
OpenAI's smartest general-purpose frontier model for professional work
Claude Opus 4.8
1M ctxAnthropic
Anthropic's top general-availability workhorse for complex agentic coding and enterprise work.
Claude Sonnet 4.6
1M ctxAnthropic
Anthropic's most capable Sonnet-class model of early 2026, now superseded by Sonnet 5.
Claude Opus 4.7
1M ctxAnthropic
The April 2026 Opus flagship - top-tier coding and vision, now superseded by Opus 4.8.
Claude Haiku 4.5
200K ctxAnthropic
Anthropic's fastest, cheapest model with near-frontier intelligence.
GPT-5
400,000 tokens (128,000 max output) ctxOpenAI
OpenAI's landmark August 2025 flagship: strong reasoning at a low price
GPT-4o
128,000 tokens (16,384 max output) ctxOpenAI
OpenAI's versatile, fast multimodal workhorse (text + image)
o1
200,000 tokens (100,000 max output) ctxOpenAI
OpenAI's first-generation deep-reasoning model that thinks before answering
Gemini 3.5
1,048,576 tokens (Gemini 3.5 Flash; Pro variant not yet released) ctxGoogle DeepMind
Google's frontier model for agents and coding, made fast and cheap.
Gemini 2.5 Pro
1,048,576 tokens (1M) input; up to 65K output ctxGoogle DeepMind
Google's advanced thinking model for complex reasoning, coding, and long context.
Gemini 2.5 Flash
1,048,576 tokens (1M) input; up to 65,535 output ctxGoogle DeepMind
Google's price-performance workhorse with thinking and a 1M-token context.
Amazon Nova Pro
300K tokens ctxAmazon Web Services
Amazon's balanced multimodal Bedrock model for text, image, and video at scale.
Mistral Large
128000 ctxMistral AI
Mistral's state-of-the-art, open-weight, general-purpose multimodal flagship.
Command R+
128K tokens ctxCohere
Cohere's RAG- and tool-use-optimized model, still live but superseded by Command A.
Claude Fable 5
1M ctxAnthropic
Anthropic's most capable widely released model - frontier intelligence for long-running agents.
North Mini Code
256K tokens ctxCohere
Cohere's first open-weight agentic coding model - 30B MoE, 3B active, runs on one H100.
GPT-5.4
1,050,000 tokens (128,000 max output) ctxOpenAI
Capable, cost-efficient predecessor to GPT-5.5 with a 1M+ context window
GPT-5.5 Instant
1050000 ctxOpenAI
The fast, default ChatGPT model tuned for low-latency responses
Mistral OCR 4
Not announced ctxMistral AI
State-of-the-art document OCR that turns PDFs and scans into structured Markdown.
GPT-5.6
1.05M ctxOpenAI
OpenAI's next-generation GPT-5.6 model family: Sol, Terra, and Luna
GPT-5.6 Sol
1050K ctxOpenAI
OpenAI's most capable and security-hardened frontier model, in limited preview
Claude Sonnet 5
1M ctxAnthropic
The best combination of speed and intelligence, at near-Opus quality for a Sonnet price.
Flint
Not announced ctxSpringboards
A divergence model designed to generate diverse, creative outputs instead of converging on predictable answers
Nano Banana 2 Lite
1M ctxThe fastest, most cost-efficient image generation model in the Nano Banana family
Muse Image
Not announced ctxMeta
Image generation model with visual reasoning capabilities built for Meta's ecosystem
GPT-Live-1
Not announced ctxOpenAI
A new generation of full-duplex voice models for natural human-AI conversation
Grok 4.5
500K ctxSpaceXAI (xAI)
Opus-class frontier model optimized for coding, agents, and knowledge work at half the cost
Robostral Navigate
Not announced ctxMistral AI
8B embodied navigation model for autonomous robot movement using single RGB camera and natural language instructions
Gemini Omni Flash
1000000 ctxCreate anything from any input with conversational video editing
Gemini 3.6 Flash
1M ctxMore token efficient and cheaper workhorse model for coding and knowledge work
Qwen 3.8 Max
983616 ctxAlibaba
2.4 trillion-parameter multimodal flagship model, claimed second only to Claude Fable 5
Gemini 3.5 Flash Cyber
1000000 ctxCost-efficient cybersecurity-focused model for vulnerability detection and patching
Claude Opus 5
1M ctxAnthropic
Thoughtful and proactive model matching Fable 5 intelligence at half the price
MAI-Cyber-1-Flash
256K ctxMicrosoft
Microsoft's first dedicated cybersecurity AI model combining efficiency and performance
Gemini Robotics ER 2
128K ctxGoogle DeepMind
High-level reasoning brain for robots enabling video understanding, task orchestration, and multi-robot collaboration
GPT-5.6-Cyber
1.05M ctxOpenAI
Purpose-trained model for advanced cybersecurity research and vulnerability testing
Palmyra X6
1000000 ctxWriter
Frontier-level performance for marketing and revenue teams with dramatically lower costs
Gemini 3.7 Flash
1M ctxMost intelligent workhorse model yet for coding and agents
Grok 4.6
500K ctxSpaceXAI
Frontier model for long-running agents, coding, and knowledge work
Open-source models
Free to download, self-host, and fine-tune.
DeepSeek V4 Flash
OpenDeepSeek
Compact 284B-param open MoE that keeps a 1M context at a fraction of Pro's cost.
DeepSeek V4
OpenDeepSeek
Open-weight 1.6T-param MoE frontier model with a 1M-token context built for agents.
DeepSeek V3
OpenDeepSeek
The open-weight 671B-param MoE that put DeepSeek on the frontier map.
Llama 4
OpenMeta
Meta's natively multimodal open MoE herd with industry-leading context length.
Qwen 3
OpenAlibaba (Qwen Team)
Alibaba's open-weight model family with switchable thinking and non-thinking modes.
Gemma 3
OpenGoogle DeepMind
Google's open, multimodal, multilingual long-context model family.
Gemma 4 12B
OpenGoogle's laptop-runnable open multimodal model with a unified encoder-free design.
GLM-5.2
OpenZhipu AI / Z.ai
Open-weights flagship model for long-horizon coding and agentic tasks with 1M-token context
Kimi K3
OpenMoonshot AI
Open frontier intelligence with 2.8T parameters and 1M context window
Laguna S 2.1
OpenPoolside
The West's most capable open-weight model for agentic coding
Antares-1B
OpenCisco
Efficient open-weight small language model for finding known vulnerabilities in codebases
Muse Glimmer
OpenMeta
Open-weight agentic AI that runs on consumer GPUs and laptops
Qwen 3.8 27B
OpenAlibaba
Apache 2.0 dense multimodal model designed for local deployment on consumer hardware
Benchmark scores sourced from official provider release posts. Prices subject to change - check provider pricing pages for current rates.