AI Models

AI Models Compared

Every major frontier model, grouped by provider - pricing, context window, benchmark scores, and what each one is actually best at. Continuously web-verified; last checked 2026-10-01.

Recently released

Pricing at a glance

ModelStrong atContextInput / 1MOutput / 1MType
OpenAI (15 models)
GPT-6.1 SolNewLong context1,050,000$2.00$10.00Closed
GPT-6 SolNewCost efficiency1.05M$2.00 (standard), $0.10 (cached input)$10.00Closed
GPT-Image-2.5 FlareNewMultimodal400000830Closed
GPT-6 AstraNewCoding1.05M$10.00$50.00Closed
AstraNewReasoning10500001050Closed
GPT-Live-1NewSpeed128000Not applicableNot applicableClosed
GPT-5.6-CyberNewPreviewCoding1.05M$12.50$75.00Closed
GPT-5.6Coding1.05M530Closed
GPT-5.6 SolPreviewCoding1050K$5.00$30.00Closed
GPT-5.5 InstantSpeed10500005.0030.00Closed
GPT-5.5Coding1,050,000 tokens (128,000 max output)$5.00$30.00Closed
GPT-5.4Long context1,050,000 tokens (128,000 max output)$2.50$15.00Closed
GPT-5Math400,000 tokens (128,000 max output)$1.25$10.00Closed
o1Reasoning200,000 tokens (100,000 max output)$15.00$60.00Closed
GPT-4oSpeed128,000 tokens (16,384 max output)$2.50$10.00Closed
Anthropic (11 models)
Claude Sonnet 5.5NewSpeed1M$2.00$10.00Closed
Claude Opus 5.5NewCoding1M$4.00$20.00Closed
Claude Mythos 5.1NewPreviewCoding1M$10.00$50.00Closed
Claude Fable 5.1New-1M$10.00$50.00Closed
Claude Opus 5Long context1M$5.00$25.00Closed
Claude Sonnet 5Coding1M$3.00$15.00Closed
Claude Fable 5CodingNot changed (1M)$10.00$0.25 (cache reads only; base output price $50 unchanged)Closed
Claude Opus 4.8Coding1M$5.00$25.00Closed
Claude Opus 4.7Coding1M$5.00$25.00Closed
Claude Sonnet 4.6Long context1M$3.00$15.00Closed
Claude Haiku 4.5Speed200K$1.00$5.00Closed
Google (10 models)
Gemini 4 ArgonNewPreviewLong context1M$2.00$10.00Closed
Gemini 3.8 Flash TTSNewSpeed8192 tokens$0.50$9.00Closed
Gemini 3.8 FlashNewCoding1M$0.75$3.75Closed
Gemini 3.5 TranscribeNewMultimodal960002.5012.00Closed
Gemini 3.7 FlashNew-1M$0.75$3.75Closed
Gemini 3.6 FlashSpeed1M$1.50$7.50Closed
Gemini 3.5 Flash CyberPreviewCoding10000001.507.50Closed
Gemini Omni FlashPreviewMultimodal1048576 tokens1.5017.50Closed
Nano Banana 2 LiteSpeed1M0.251.50Closed
Gemma 4 12BCost efficiency256K tokensFreeFreeOpen
Google DeepMind (6 models)
Gemini 3.8 LiveNewSpeed128K$3.00$12.00Closed
Gemini Robotics ER 2Multimodal128K$2$10Closed
Gemini 3.5Coding1,048,576 tokens (Gemini 3.5 Flash; Pro variant not yet released)$1.50$9.00Closed
Gemini 2.5 ProLong context1,048,576 tokens (1M) input; up to 65K output$1.25$10.00Closed
Gemini 2.5 FlashCost efficiency1,048,576 tokens (1M) input; up to 65,535 output$0.30$2.50Closed
Gemma 3Cost efficiency128K tokens (32K for the 1B variant)FreeFreeOpen
Alibaba (3 models)
Qwen 3.8-Flash-NextNewCost efficiency262KFreeFreeOpen
Qwen 3.8 27BNewCoding262KFreeFreeOpen
Qwen 3.8 MaxNewLong context9836162.006.00Closed
Cohere (3 models)
Cohere Parse 5NewCost efficiency81921.50Not announcedClosed
North Mini CodeCost efficiency256K tokens00Closed
Command R+Long context128K tokens$2.50$10.00Closed
DeepSeek (3 models)
DeepSeek V4NewPreviewLong context1M tokensFreeFreeOpen
DeepSeek V4 FlashPreviewLong context1M tokensFreeFreeOpen
DeepSeek V3Cost efficiency128K tokensFreeFreeOpen
Meta (3 models)
Muse GlimmerNewCost efficiency128KFreeFreeOpen
Muse ImageCost efficiency65536$0.01/image$0.01/imageClosed
Llama 4Long contextUp to 10M tokens (Scout); ~1M tokens (Maverick)FreeFreeOpen
Mistral AI (3 models)
Robostral NavigateCost efficiencyNot announcedNot announcedNot announcedClosed
Mistral OCR 4Cost efficiency16K$4.00 per 1,000 pages (API) / $2.00 per 1,000 pages (Batch)$5.00 per 1,000 pages (Document AI annotation tier)Closed
Mistral LargeReasoning1280002.006.00Closed
Alibaba (Qwen Team) (1 model)
Qwen 3Cost efficiency128K tokens (32K for 0.6B/1.7B/4B dense variants)FreeFreeOpen
Amazon Web Services (1 model)
Amazon Nova ProMultimodal300K tokens$0.80$3.20Closed
Anonymous (Stealth) (1 model)
Ox AlphaNewPreviewLong context1,048,576 (1M tokens)$0.00$0.00Closed
Cisco (1 model)
Antares-1BCost efficiency128000FreeFreeOpen
Cognition (1 model)
SWE-2NewCoding1M3.0015.00Closed
Microsoft (1 model)
MAI-Cyber-1-FlashPreviewCoding256000Not announcedNot announcedClosed
Moonshot AI (1 model)
Kimi K3Coding1M$3.00$15.00Open
Poolside (1 model)
Laguna S 2.1Coding1MNot announcedNot announcedOpen
Salesforce (1 model)
Salesforce KoaNewPreviewReasoning1000000Not announcedNot announcedOpen
SpaceXAI (1 model)
Grok 4.6NewCost efficiency500K$2.00$6.00Closed
SpaceXAI (xAI) (1 model)
Grok 4.5Cost efficiency500K$2.00$6.00Closed
Springboards (1 model)
FlintPreviewSpeedNot announcedNot announcedNot announcedClosed
Thomson Reuters (1 model)
ThomsonNewCost efficiency262000Not announcedNot announcedClosed
TypeSafe AI (1 model)
JevNewPreviewCost efficiency32K$0.042FreeClosed
Writer (1 model)
Palmyra X6NewCost efficiency1000000$2.00$8.00Closed
xAI (SpaceXAI) (1 model)
Grok 4.7NewCoding500K$2.00$6.00Closed
Z.ai (Zhipu AI) (1 model)
GLM-5.3-FlashNewLong context1M$0.075$0.25Open
Zhipu AI / Z.ai (1 model)
GLM-5.2Long context1M$1.40$4.40Open

Prices in USD. Open-source models are free to self-host; API pricing varies by provider.

OpenAI15 models

GPT-6.1 SolNew

1,050,000 ctx

2026-09

Near-Astra intelligence for a fifth of the price

Agentic coding workflows and computer use automationLong-running software engineering tasksDocument-heavy professional work and analysis
Full details →

GPT-6 SolNew

1.05M ctx

2026-09

Cost-efficient mid-tier model for complex professional work and coding

Complex coding and agentic workflowsBusiness workflow automationSoftware engineering and code review
Full details →

GPT-Image-2.5 FlareNew

400000 ctx

2026-09

Faster, sharper, smarter image generation

Fast image generation workflowsProduction-scale image creationIterative creative refinement
Full details →

GPT-6 AstraNew

1.05M ctx

2026-09

The world's most intelligent and aligned model

Complex reasoning and problem-solvingAgentic workflows and autonomous workScientific research and discovery
Full details →

AstraNew

1050000 ctx

2026-09

OpenAI's next major model with critical cybersecurity capabilities

Advanced mathematical research and proofsDefensive cybersecurity applicationsLong-horizon agentic tasks requiring persistent planning
Full details →

GPT-Live-1New

128000 ctx

2026-09

A new generation of full-duplex voice models for natural human-AI conversation

Hands-free assistance with cooking, directions, and daily tasksLanguage practice with natural back-and-forth dialogueLive translation during conversations
Full details →

GPT-5.6-CyberNew

1.05M ctx

2026-08

Purpose-trained model for advanced cybersecurity research and vulnerability testing

Vulnerability research and exploitation testingPenetration testing and red-team exercisesExploit validation on authorized systems
Full details →

GPT-5.6

1.05M ctx

2026-07

OpenAI's next-generation GPT-5.6 model family: Sol, Terra, and Luna

Teams that want one generation with multiple cost/quality tiersAgentic coding, knowledge work, and research at scaleEarly adopters and approved preview partners
Full details →

GPT-5.6 Sol

1050K ctx

2026-06

OpenAI's most capable and security-hardened frontier model, in limited preview

Frontier agentic coding and computer useComplex multi-step tasks via subagent orchestrationSecurity-sensitive and scientific research workloads
Full details →

GPT-5.5 Instant

1050000 ctx

2026-05

The fast, default ChatGPT model tuned for low-latency responses

Fast, everyday ChatGPT interactionsLatency-sensitive assistant and chat useQuick drafting and summarization
Full details →

GPT-5.5

1,050,000 tokens (128,000 max output) ctx

2026-04

OpenAI's smartest general-purpose frontier model for professional work

Agentic coding and multi-tool automationLarge-context document and codebase analysisGeneral-purpose frontier reasoning and knowledge work
Full details →

GPT-5.4

1,050,000 tokens (128,000 max output) ctx

2026-03

Capable, cost-efficient predecessor to GPT-5.5 with a 1M+ context window

Cost-efficient large-context workGeneral coding and reasoning at scaleProduction workloads migrating within the GPT-5 generation
Full details →

GPT-5

400,000 tokens (128,000 max output) ctx

2025-08

OpenAI's landmark August 2025 flagship: strong reasoning at a low price

Budget-friendly frontier reasoning and codingMath and STEM problem solvingExisting GPT-5-based production systems
Full details →

o1

200,000 tokens (100,000 max output) ctx

2024-12

OpenAI's first-generation deep-reasoning model that thinks before answering

Hard math and science reasoningCompetitive and algorithmic programmingDeliberate multi-step problem solving
Full details →

GPT-4o

128,000 tokens (16,384 max output) ctx

2024-05

OpenAI's versatile, fast multimodal workhorse (text + image)

Fast, low-cost everyday assistant tasksMultimodal (image + text) understandingHigh-volume production workloads
Full details →

Anthropic11 models

Claude Sonnet 5.5New

1M ctx

2026-09

Faster, cheaper mid-tier model for everyday tasks and agentic work

Everyday tasks and workflowsAgentic coding and automationDocument, slide, and spreadsheet creation
Full details →

Claude Opus 5.5New

1M ctx

2026-09

The strongest-performing model we've tested to date

Complex coding and agentic tasksLong-horizon reasoningFinancial analysis and compliance work
Full details →

Claude Mythos 5.1New

1M ctx

2026-09

Anthropic's frontier-tier model designed for advanced cybersecurity and critical infrastructure work

Critical software infrastructure securityAdvanced vulnerability assessmentComplex reasoning and long-context tasks
Full details →

Claude Fable 5.1New

1M ctx

2026-09

Point release of Anthropic's flagship, tuned for long-horizon agentic work with a large effective price cut.

Long-running agentic pipelines where cache-read cost dominatesAutomated scientific research and computational biology workflowsDefensive security work (via the Cyber Verification Program for Mythos 5.1)
Full details →

Claude Opus 5

1M ctx

2026-07

Thoughtful and proactive model matching Fable 5 intelligence at half the price

coding taskslong-running agentsprofessional knowledge work
Full details →

Claude Sonnet 5

1M ctx

2026-06

The best combination of speed and intelligence, at near-Opus quality for a Sonnet price.

Agentic coding assistants and autonomous dev agentsHigh-volume production agent loops needing strong reasoning per dollarLarge-codebase and long-document analysis with 1M context
Full details →

Claude Fable 5

Not changed (1M) ctx

2026-06

Anthropic's most capable widely released model - frontier intelligence for long-running agents.

Long-running autonomous agents on complex, high-value tasksFrontier software engineering and hard multi-service implementationScientific research and advanced analytics
Full details →

Claude Opus 4.8

1M ctx

2026-05

Anthropic's top general-availability workhorse for complex agentic coding and enterprise work.

Complex, long-horizon agentic coding and autonomous engineeringEnterprise knowledge work and professional-grade analysisHigh-stakes tasks where honesty and flaw-catching matter
Full details →

Claude Opus 4.7

1M ctx

2026-04

The April 2026 Opus flagship - top-tier coding and vision, now superseded by Opus 4.8.

Vision-heavy agentic and document-understanding workloadsTop-tier coding on established Opus 4.7 pipelinesLarge-codebase engineering with 1M context
Full details →

Claude Sonnet 4.6

1M ctx

2026-02

Anthropic's most capable Sonnet-class model of early 2026, now superseded by Sonnet 5.

Established mid-tier production workloads pinned to a stable modelLarge-context document and codebase analysisAgentic coding at a moderate price point
Full details →

Claude Haiku 4.5

200K ctx

2025-10

Anthropic's fastest, cheapest model with near-frontier intelligence.

High-throughput, low-latency production workloadsParallelized sub-agents and multi-agent worker rolesCheap classification, extraction, and summarization at scale
Full details →

Google10 models

Gemini 4 ArgonNew

1M ctx

2026-09

Google's frontier model for real-world coding, enterprise knowledge work, and cybersecurity defense

Long-form software engineering tasksCybersecurity vulnerability analysisEnterprise knowledge work (legal, finance)
Full details →

Gemini 3.8 Flash TTSNew

8192 tokens ctx

2026-09

Advanced text-to-speech with natural language voice design and line-by-line performance direction

Immersive audiobooksGame character voicesPodcast narration
Full details →

Gemini 3.8 FlashNew

1M ctx

2026-09

Most intelligent workhorse model with enhanced reasoning and coding

Software engineering tasksAutonomous agentsComplex enterprise analysis
Full details →

Gemini 3.5 TranscribeNew

96000 ctx

2026-08

The most precise speech-to-text model yet

Real-time voice agents and applicationsPost-call analytics and meeting transcriptionMultilingual voice interactions
Full details →

Gemini 3.7 FlashNew

1M ctx

2026-08

Most intelligent workhorse model yet for coding and agents

Software engineering and code generationAutonomous agent workflows and business automationDocument-heavy knowledge work
Full details →

Gemini 3.6 Flash

1M ctx

2026-07

More token efficient and cheaper workhorse model for coding and knowledge work

Cost-sensitive production AI agentsAgentic workflows requiring efficiencyMulti-step coding tasks
Full details →

Gemini 3.5 Flash Cyber

1000000 ctx

2026-07

Cost-efficient cybersecurity-focused model for vulnerability detection and patching

Automated vulnerability detection at scaleCost-sensitive security scanning workflowsOrganizations needing efficient code security analysis
Full details →

Gemini Omni Flash

1048576 tokens ctx

2026-06

Create anything from any input with conversational video editing

Content creators needing iterative video editingBusinesses creating marketing and training videosUsers wanting to generate videos with custom avatars
Full details →

Nano Banana 2 Lite

1M ctx

2026-06

The fastest, most cost-efficient image generation model in the Nano Banana family

High-volume image generationRapid ideation and A/B testingIterative content creation
Full details →

Gemma 4 12B

256K tokens ctx

2026-06

Open

Google's laptop-runnable open multimodal model with a unified encoder-free design.

Self-hosted multimodal assistantsOn-device / single-GPU deploymentPrivacy-sensitive applications
Full details →

Google DeepMind6 models

Gemini 3.8 LiveNew

128K ctx

2026-09

Most advanced live dialogue model for natural, real-time voice conversations

Production voice agentsReal-time conversational AICustomer service automation
Full details →

Gemini Robotics ER 2

128K ctx

2026-07

High-level reasoning brain for robots enabling video understanding, task orchestration, and multi-robot collaboration

Complex multi-step robotic tasksMulti-robot coordination and collaborationTasks requiring progress monitoring and temporal understanding
Full details →

Gemini 3.5

1,048,576 tokens (Gemini 3.5 Flash; Pro variant not yet released) ctx

2026-05

Google's frontier model for agents and coding, made fast and cheap.

Autonomous coding agentsHigh-volume production apps needing frontier qualityMultimodal document and chart understanding
Full details →

Gemini 2.5 Pro

1,048,576 tokens (1M) input; up to 65K output ctx

2025-06

Google's advanced thinking model for complex reasoning, coding, and long context.

Complex reasoning and STEM problem-solvingLong-context document and codebase analysisHigh-quality multimodal understanding
Full details →

Gemini 2.5 Flash

1,048,576 tokens (1M) input; up to 65,535 output ctx

2025-06

Google's price-performance workhorse with thinking and a 1M-token context.

High-volume production workloadsCost-sensitive chat and RAG appsFast multimodal processing
Full details →

Gemma 3

128K tokens (32K for the 1B variant) ctx

2025-03

Open

Google's open, multimodal, multilingual long-context model family.

Local and on-prem multimodal assistantsMultilingual applicationsFine-tuning on custom data
Full details →

Alibaba3 models

Cohere3 models

DeepSeek3 models

Meta3 models

Mistral AI3 models

Alibaba (Qwen Team)1 model

Amazon Web Services1 model

Anonymous (Stealth)1 model

Cisco1 model

Cognition1 model

Microsoft1 model

Moonshot AI1 model

Poolside1 model

Salesforce1 model

SpaceXAI1 model

SpaceXAI (xAI)1 model

Springboards1 model

Thomson Reuters1 model

TypeSafe AI1 model

Writer1 model

xAI (SpaceXAI)1 model

Z.ai (Zhipu AI)1 model

Zhipu AI / Z.ai1 model

Benchmark scores sourced from official provider release posts. Prices subject to change - check provider pricing pages for current rates.