announcementsai-trends

GPT-6 Astra lands: same price as Claude's flagship, wilder benchmarks

OpenAI's new flagship posts near-perfect scores on ARC-AGI-3 and FrontierMath, rolls out without a free tier, and costs exactly what Anthropic charges for Fable 5.1. The frontier now has a sticker price: $10 in, $50 out.

September 4, 2026

GPT-6 Astra lands: same price as Claude's flagship, wilder benchmarks

OpenAI released GPT-6 Astra on September 3, its new flagship for reasoning, software engineering, computer use, scientific research, and cybersecurity work. The rollout started inside OpenAI's Daybreak cybersecurity program, with API access, Microsoft Azure, AWS Bedrock, and the ChatGPT Plus, Pro, Business, and Enterprise plans following over the coming days. There is no free-tier access. We have added it to our model database.

The benchmark everyone is quoting, and what it actually says

The number driving the AGI chatter is 99.9% on ARC-AGI-3, the benchmark built around novel turn-based environments where an agent has to explore, infer the goal, and plan without instructions. The ARC Prize team's own writeup deserves a closer read than the headline, because the details are stranger than the score.

Two results stand out. First, the score splits by harness: 62.7% under the standard harness, 99.9% under the provider-adapter harness. Second, efficiency, which is what our cover chart shows: Astra completed levels using fewer actions than the human baseline on 96.0% of levels, averaging 51.7% fewer actions per level. It also invented its own compact algebraic notation to track game state. Roughly 500 members of the general public set the human baseline, at an average cost of $12.78 per attempted game.

"Saturating the benchmark would not represent proof of achieving AGI." - ARC Prize team, on their own benchmark

That caveat is from the people who built the test. ARC-AGI-3 has deterministic, closed-ended mechanics; real work does not. The same tempering applies to the other near-perfect scores OpenAI reports: 98% on FrontierMath Tier 4 and 100% on ExploitBench. When a model saturates a benchmark, the honest conclusion is that the benchmark is finished, not that the work is.

The price is the tell

GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens. If that sounds familiar, it is because Claude Fable 5.1, released two days earlier, costs exactly the same. Two frontier flagships, same week, identical list price.

GPT-6 AstraClaude Fable 5.1
Input / 1M tokens$10.00$10.00
Output / 1M tokens$50.00$50.00
Context window1.05M1M
ReleasedSeptember 3, 2026September 1, 2026
Speed optionFast mode: 2x speed at 2x priceStandard
Effective cost leverNone announced at launchCache reads $0.25/1M; 25-45% lower typical workload cost

This is what a settled market segment looks like. The frontier tier now has a sticker price, and the competition has moved to effective cost: Anthropic is discounting through token efficiency and cheap cache reads, OpenAI is upselling speed at double rate and a higher-effort Astra Pro variant for Pro, Business, and Enterprise plans. For anyone running agents at scale, the invoice will differ far more by workload shape than by list price. Our AI Pricing Index tracks both, and every price change since May.

What it means if you use ChatGPT

Practical facts for subscribers, in rollout order:

  1. Plus, Pro, Business, and Enterprise users get GPT-6 Astra inside existing plan allowances over the coming days, with the option to buy credits beyond them.
  2. Pro, Business, and Enterprise plans additionally get GPT-6 Astra Pro, a higher-effort variant for harder tasks.
  3. Free-tier users get nothing this time. Every earlier flagship eventually trickled down; OpenAI has not said that here.

For most ChatGPT subscribers the honest advice is unchanged from every flagship launch: the difference shows up on long, hard, multi-step tasks, not in everyday chat. If your use is drafting, summarizing, and quick questions, the model picker matters less than the workflow around it. If you are deciding between ecosystems, our ChatGPT vs Claude comparison covers the ground that benchmarks do not.

The pattern worth watching

Astra's biggest claimed gains are in agentic work: computer use, autonomous task execution, cybersecurity. Fable 5.1's biggest gains two days earlier were also agentic: scientific research pipelines and long-horizon tool use. Both companies shipped their September releases aimed at the same target, and it is not chat. The frontier labs are now competing on how long a model can work unattended, and pricing that work is where the real fight is. Watch the model index; September is not done.

Some links in this article are affiliate links. Learn more.