AI Models

AI Models Compared

Every major frontier model - pricing, context window, benchmark scores, and what each one is actually best at. Continuously web-verified; last checked 2026-08-20.

Pricing at a glance

ModelProviderStrong atContextInput / 1MOutput / 1MType
DeepSeek V4 FlashPreviewDeepSeekLong context1M tokensFreeFreeOpen
DeepSeek V4PreviewDeepSeekLong context1M tokensFreeFreeOpen
GPT-5.5OpenAICoding1,050,000 tokens (128,000 max output)$5.00$30.00Closed
Claude Opus 4.8AnthropicCoding1M$5.00$25.00Closed
Claude Sonnet 4.6AnthropicLong context1M$3.00$15.00Closed
Claude Opus 4.7AnthropicCoding1M$5.00$25.00Closed
Claude Haiku 4.5AnthropicSpeed200K$1.00$5.00Closed
GPT-5OpenAIMath400,000 tokens (128,000 max output)$1.25$10.00Closed
GPT-4oOpenAISpeed128,000 tokens (16,384 max output)$2.50$10.00Closed
o1OpenAIReasoning200,000 tokens (100,000 max output)$15.00$60.00Closed
Gemini 3.5Google DeepMindCoding1,048,576 tokens (Gemini 3.5 Flash; Pro variant not yet released)$1.50$9.00Closed
Gemini 2.5 ProGoogle DeepMindLong context1,048,576 tokens (1M) input; up to 65K output$1.25$10.00Closed
Gemini 2.5 FlashGoogle DeepMindCost efficiency1,048,576 tokens (1M) input; up to 65,535 output$0.30$2.50Closed
Amazon Nova ProAmazon Web ServicesMultimodal300K tokens$0.80$3.20Closed
DeepSeek V3DeepSeekCost efficiency128K tokensFreeFreeOpen
Llama 4MetaLong contextUp to 10M tokens (Scout); ~1M tokens (Maverick)FreeFreeOpen
Qwen 3Alibaba (Qwen Team)Cost efficiency128K tokens (32K for 0.6B/1.7B/4B dense variants)FreeFreeOpen
Mistral LargeMistral AIReasoning1280002.006.00Closed
Gemma 3Google DeepMindCost efficiency128K tokens (32K for the 1B variant)FreeFreeOpen
Command R+CohereLong context128K tokens$2.50$10.00Closed
Gemma 4 12BGoogleCost efficiency256K tokensFreeFreeOpen
Claude Fable 5AnthropicCoding1M$10.00$50.00Closed
North Mini CodeCohereCost efficiency256K tokens00Closed
GPT-5.4OpenAILong context1,050,000 tokens (128,000 max output)$2.50$15.00Closed
GPT-5.5 InstantOpenAISpeed10500005.0030.00Closed
Mistral OCR 4Mistral AICost efficiencyNot announced$4.00 per 1,000 pages (API) / $2.00 per 1,000 pages (Batch)$5.00 per 1,000 pages (Document AI annotation tier)Closed
GPT-5.6OpenAICoding1.05M530Closed
GPT-5.6 SolPreviewOpenAICoding1050K$5.00$30.00Closed
Claude Sonnet 5AnthropicCoding1M$3.00$15.00Closed
FlintPreviewSpringboardsSpeedNot announcedNot announcedNot announcedClosed
Nano Banana 2 LiteGoogleSpeed1M0.251.50Closed
GLM-5.2Zhipu AI / Z.aiLong context1M$1.40$4.40Open
Muse ImageMetaCost efficiencyNot announcedNot announcedNot announcedClosed
GPT-Live-1OpenAISpeedNot announcedNot announcedNot announcedClosed
Grok 4.5SpaceXAI (xAI)Cost efficiency500K$2.00$6.00Closed
Robostral NavigateMistral AICost efficiencyNot announcedNot announcedNot announcedClosed
Gemini Omni FlashPreviewGoogleMultimodal10000001.5017.50Closed
Kimi K3Moonshot AICoding1M$3.00$15.00Open
Gemini 3.6 FlashGoogleSpeed1M$1.50$7.50Closed
Laguna S 2.1PoolsideCoding1MNot announcedNot announcedOpen
Qwen 3.8 MaxAlibabaLong context9836162.006.00Closed
Gemini 3.5 Flash CyberPreviewGoogleCoding10000001.507.50Closed
Antares-1BCiscoCost efficiency128000FreeFreeOpen
Claude Opus 5AnthropicLong context1M$5.00$25.00Closed
MAI-Cyber-1-FlashPreviewMicrosoftCoding256KNot announcedNot announcedClosed
Gemini Robotics ER 2Google DeepMindMultimodal128K$2$10Closed
GPT-5.6-CyberPreviewOpenAICoding1.05M$12.50$75.00Closed
Muse GlimmerMetaCost efficiency128KFreeFreeOpen
Palmyra X6WriterCost efficiency1000000$2.00$8.00Closed
Gemini 3.7 FlashGoogleLong context1M$0.75$3.75Closed
Grok 4.6SpaceXAICost efficiency500K$2.00$6.00Closed
Qwen 3.8 27BAlibabaCoding262KFreeFreeOpen

Prices in USD. Open-source models are free to self-host; API pricing varies by provider.

Closed-source models

GPT-5.5

1,050,000 tokens (128,000 max output) ctx

OpenAI

OpenAI's smartest general-purpose frontier model for professional work

Agentic coding and multi-tool automationLarge-context document and codebase analysisGeneral-purpose frontier reasoning and knowledge work
Full details →

Claude Opus 4.8

1M ctx

Anthropic

Anthropic's top general-availability workhorse for complex agentic coding and enterprise work.

Complex, long-horizon agentic coding and autonomous engineeringEnterprise knowledge work and professional-grade analysisHigh-stakes tasks where honesty and flaw-catching matter
Full details →

Claude Sonnet 4.6

1M ctx

Anthropic

Anthropic's most capable Sonnet-class model of early 2026, now superseded by Sonnet 5.

Established mid-tier production workloads pinned to a stable modelLarge-context document and codebase analysisAgentic coding at a moderate price point
Full details →

Claude Opus 4.7

1M ctx

Anthropic

The April 2026 Opus flagship - top-tier coding and vision, now superseded by Opus 4.8.

Vision-heavy agentic and document-understanding workloadsTop-tier coding on established Opus 4.7 pipelinesLarge-codebase engineering with 1M context
Full details →

Claude Haiku 4.5

200K ctx

Anthropic

Anthropic's fastest, cheapest model with near-frontier intelligence.

High-throughput, low-latency production workloadsParallelized sub-agents and multi-agent worker rolesCheap classification, extraction, and summarization at scale
Full details →

GPT-5

400,000 tokens (128,000 max output) ctx

OpenAI

OpenAI's landmark August 2025 flagship: strong reasoning at a low price

Budget-friendly frontier reasoning and codingMath and STEM problem solvingExisting GPT-5-based production systems
Full details →

GPT-4o

128,000 tokens (16,384 max output) ctx

OpenAI

OpenAI's versatile, fast multimodal workhorse (text + image)

Fast, low-cost everyday assistant tasksMultimodal (image + text) understandingHigh-volume production workloads
Full details →

o1

200,000 tokens (100,000 max output) ctx

OpenAI

OpenAI's first-generation deep-reasoning model that thinks before answering

Hard math and science reasoningCompetitive and algorithmic programmingDeliberate multi-step problem solving
Full details →

Gemini 3.5

1,048,576 tokens (Gemini 3.5 Flash; Pro variant not yet released) ctx

Google DeepMind

Google's frontier model for agents and coding, made fast and cheap.

Autonomous coding agentsHigh-volume production apps needing frontier qualityMultimodal document and chart understanding
Full details →

Gemini 2.5 Pro

1,048,576 tokens (1M) input; up to 65K output ctx

Google DeepMind

Google's advanced thinking model for complex reasoning, coding, and long context.

Complex reasoning and STEM problem-solvingLong-context document and codebase analysisHigh-quality multimodal understanding
Full details →

Gemini 2.5 Flash

1,048,576 tokens (1M) input; up to 65,535 output ctx

Google DeepMind

Google's price-performance workhorse with thinking and a 1M-token context.

High-volume production workloadsCost-sensitive chat and RAG appsFast multimodal processing
Full details →

Amazon Nova Pro

300K tokens ctx

Amazon Web Services

Amazon's balanced multimodal Bedrock model for text, image, and video at scale.

AWS/Bedrock-native applicationsMultimodal text/image/video tasksCost-balanced general workloads
Full details →

Mistral Large

128000 ctx

Mistral AI

Mistral's state-of-the-art, open-weight, general-purpose multimodal flagship.

Multilingual reasoning and generationStructured/JSON output and codingOpen-weight flagship deployments
Full details →

Command R+

128K tokens ctx

Cohere

Cohere's RAG- and tool-use-optimized model, still live but superseded by Command A.

Retrieval-augmented generation with citationsMulti-step tool-use / agentsMultilingual enterprise assistants
Full details →

Claude Fable 5

1M ctx

Anthropic

Anthropic's most capable widely released model - frontier intelligence for long-running agents.

Long-running autonomous agents on complex, high-value tasksFrontier software engineering and hard multi-service implementationScientific research and advanced analytics
Full details →

North Mini Code

256K tokens ctx

Cohere

Cohere's first open-weight agentic coding model - 30B MoE, 3B active, runs on one H100.

Agentic software engineeringTerminal/tool-driven coding agentsSelf-hosted sovereign coding on one GPU
Full details →

GPT-5.4

1,050,000 tokens (128,000 max output) ctx

OpenAI

Capable, cost-efficient predecessor to GPT-5.5 with a 1M+ context window

Cost-efficient large-context workGeneral coding and reasoning at scaleProduction workloads migrating within the GPT-5 generation
Full details →

GPT-5.5 Instant

1050000 ctx

OpenAI

The fast, default ChatGPT model tuned for low-latency responses

Fast, everyday ChatGPT interactionsLatency-sensitive assistant and chat useQuick drafting and summarization
Full details →

Mistral OCR 4

Not announced ctx

Mistral AI

State-of-the-art document OCR that turns PDFs and scans into structured Markdown.

Document OCR and parsingPDF/scan to structured Markdown or JSONMultilingual document intelligence
Full details →

GPT-5.6

1.05M ctx

OpenAI

OpenAI's next-generation GPT-5.6 model family: Sol, Terra, and Luna

Teams that want one generation with multiple cost/quality tiersAgentic coding, knowledge work, and research at scaleEarly adopters and approved preview partners
Full details →

GPT-5.6 Sol

1050K ctx

OpenAI

OpenAI's most capable and security-hardened frontier model, in limited preview

Frontier agentic coding and computer useComplex multi-step tasks via subagent orchestrationSecurity-sensitive and scientific research workloads
Full details →

Claude Sonnet 5

1M ctx

Anthropic

The best combination of speed and intelligence, at near-Opus quality for a Sonnet price.

Agentic coding assistants and autonomous dev agentsHigh-volume production agent loops needing strong reasoning per dollarLarge-codebase and long-document analysis with 1M context
Full details →

Flint

Not announced ctx

Springboards

A divergence model designed to generate diverse, creative outputs instead of converging on predictable answers

Advertising and marketing ideationCreative brainstormingStrategic exploration
Full details →

Nano Banana 2 Lite

1M ctx

Google

The fastest, most cost-efficient image generation model in the Nano Banana family

High-volume image generationRapid ideation and A/B testingIterative content creation
Full details →

Muse Image

Not announced ctx

Meta

Image generation model with visual reasoning capabilities built for Meta's ecosystem

Casual Instagram and WhatsApp users generating images in-appAdvertisers creating multiple ad variations automaticallyUsers wanting quick image edits without leaving social apps
Full details →

GPT-Live-1

Not announced ctx

OpenAI

A new generation of full-duplex voice models for natural human-AI conversation

Hands-free assistance with cooking, directions, and daily tasksLanguage practice with natural back-and-forth dialogueLive translation during conversations
Full details →

Grok 4.5

500K ctx

SpaceXAI (xAI)

Opus-class frontier model optimized for coding, agents, and knowledge work at half the cost

Cost-conscious teams running high-volume agentic workloadsCoding and software engineering tasksEnterprise knowledge work and reasoning
Full details →

Robostral Navigate

Not announced ctx

Mistral AI

8B embodied navigation model for autonomous robot movement using single RGB camera and natural language instructions

Indoor autonomous navigation in offices, warehouses, and commercial buildingsManufacturing, logistics, and delivery applicationsResource-constrained deployments requiring low hardware costs
Full details →

Gemini Omni Flash

1000000 ctx

Google

Create anything from any input with conversational video editing

Content creators needing iterative video editingBusinesses creating marketing and training videosUsers wanting to generate videos with custom avatars
Full details →

Gemini 3.6 Flash

1M ctx

Google

More token efficient and cheaper workhorse model for coding and knowledge work

Cost-sensitive production AI agentsAgentic workflows requiring efficiencyMulti-step coding tasks
Full details →

Qwen 3.8 Max

983616 ctx

Alibaba

2.4 trillion-parameter multimodal flagship model, claimed second only to Claude Fable 5

Coding and full-stack developmentData analysis workflowsOffice/productivity task automation
Full details →

Gemini 3.5 Flash Cyber

1000000 ctx

Google

Cost-efficient cybersecurity-focused model for vulnerability detection and patching

Automated vulnerability detection at scaleCost-sensitive security scanning workflowsOrganizations needing efficient code security analysis
Full details →

Claude Opus 5

1M ctx

Anthropic

Thoughtful and proactive model matching Fable 5 intelligence at half the price

coding taskslong-running agentsprofessional knowledge work
Full details →

MAI-Cyber-1-Flash

256K ctx

Microsoft

Microsoft's first dedicated cybersecurity AI model combining efficiency and performance

Enterprise vulnerability discovery and remediationCost-efficient security scanning at scaleOrganizations seeking cybersecurity AI without government restrictions
Full details →

Gemini Robotics ER 2

128K ctx

Google DeepMind

High-level reasoning brain for robots enabling video understanding, task orchestration, and multi-robot collaboration

Complex multi-step robotic tasksMulti-robot coordination and collaborationTasks requiring progress monitoring and temporal understanding
Full details →

GPT-5.6-Cyber

1.05M ctx

OpenAI

Purpose-trained model for advanced cybersecurity research and vulnerability testing

Vulnerability research and exploitation testingPenetration testing and red-team exercisesExploit validation on authorized systems
Full details →

Palmyra X6

1000000 ctx

Writer

Frontier-level performance for marketing and revenue teams with dramatically lower costs

High-volume marketing and sales AI agentsCost-sensitive enterprise deployments at scaleLong-running agentic workflows
Full details →

Gemini 3.7 Flash

1M ctx

Google

Most intelligent workhorse model yet for coding and agents

Software engineering and code generationAutonomous agent workflows and business automationDocument-heavy knowledge work
Full details →

Grok 4.6

500K ctx

SpaceXAI

Frontier model for long-running agents, coding, and knowledge work

Long-running agentic tasksComplex coding and software engineeringKnowledge work and reasoning
Full details →

Open-source models

Free to download, self-host, and fine-tune.

DeepSeek V4 Flash

Open

DeepSeek

Compact 284B-param open MoE that keeps a 1M context at a fraction of Pro's cost.

High-volume, cost-sensitive inferenceLong-context tasks on a budgetSelf-hosting on modest hardware
Full details →

DeepSeek V4

Open

DeepSeek

Open-weight 1.6T-param MoE frontier model with a 1M-token context built for agents.

Self-hosted frontier reasoning and codingLong-context agentic workflowsCost-sensitive high-volume inference
Full details →

DeepSeek V3

Open

DeepSeek

The open-weight 671B-param MoE that put DeepSeek on the frontier map.

Cost-efficient open-weight general useSelf-hosting and fine-tuningCoding and instruction-following
Full details →

Llama 4

Open

Meta

Meta's natively multimodal open MoE herd with industry-leading context length.

Extremely long-context open-model applicationsSelf-hosted multimodal deploymentsEfficient MoE inference
Full details →

Qwen 3

Open

Alibaba (Qwen Team)

Alibaba's open-weight model family with switchable thinking and non-thinking modes.

Self-hosted reasoning and coding assistantsApplications needing switchable reasoning depthMultilingual open-model deployments
Full details →

Gemma 3

Open

Google DeepMind

Google's open, multimodal, multilingual long-context model family.

Local and on-prem multimodal assistantsMultilingual applicationsFine-tuning on custom data
Full details →

Gemma 4 12B

Open

Google

Google's laptop-runnable open multimodal model with a unified encoder-free design.

Self-hosted multimodal assistantsOn-device / single-GPU deploymentPrivacy-sensitive applications
Full details →

GLM-5.2

Open

Zhipu AI / Z.ai

Open-weights flagship model for long-horizon coding and agentic tasks with 1M-token context

Long-horizon coding and agentic automation tasksRepository-scale software engineering with full codebase contextMulti-file refactoring and complex engineering workflows
Full details →

Kimi K3

Open

Moonshot AI

Open frontier intelligence with 2.8T parameters and 1M context window

Coding and software engineeringLong-document analysis and researchComplex reasoning and agentic workflows
Full details →

Laguna S 2.1

Open

Poolside

The West's most capable open-weight model for agentic coding

Agentic coding tasksOrganizations requiring local deploymentCost-efficient coding model deployment
Full details →

Antares-1B

Open

Cisco

Efficient open-weight small language model for finding known vulnerabilities in codebases

Security teams with limited resourcesOrganizations requiring local code analysisContinuous vulnerability scanning with cost constraints
Full details →

Muse Glimmer

Open

Meta

Open-weight agentic AI that runs on consumer GPUs and laptops

Local AI agentsCoding agentsTool use and function calling
Full details →

Qwen 3.8 27B

Open

Alibaba

Apache 2.0 dense multimodal model designed for local deployment on consumer hardware

Local AI deployment on edge devices and workstationsSoftware engineering and coding tasksDocument analysis and office productivity
Full details →

Benchmark scores sourced from official provider release posts. Prices subject to change - check provider pricing pages for current rates.