GPT-6.1 Sol arrives. OpenAI's new model matches Astra at a fifth the cost.
OpenAI has released GPT-6.1 Sol, a more affordable model delivering near-Astra level performance at significantly lower pricing, expanding access to high-capability AI.
September 30, 2026

You are choosing between GPT-6 Astra and GPT-6.1 Sol because your API bill and your output quality are pulling in opposite directions, and OpenAI just gave you a third option that claims to close that gap. The pitch, straight from OpenAI's own announcement, is that Sol gets "near-Astra" intelligence at roughly a fifth of the cost. That is a specific, testable claim. Here is what it actually means for a team deciding which model name to put in production.
The case for staying on Astra anyway
"Near" is the word doing all the work in that headline, and it is worth taking seriously before you touch a single API call. GPT-6 Astra runs $10.00 input and $50.00 output per 1M tokens. GPT-6.1 Sol runs $2.00 and $10.00. That is a real 5x cut, not a rounding trick. But price-per-token and price-per-completed-task are different numbers, and the gap between them is where these announcements usually fall apart.
A cheaper model that needs more tool calls, more retries, or a longer chain-of-thought to reach the same answer can end up costing more per finished task than the expensive model it was supposed to replace. This is not a hypothetical. It is the exact failure mode that shows up every time a lab ships a "distilled" or "mini" tier: strong on benchmark suites built for single-turn evaluation, weaker on the messy multi-step agentic work that dominates real production traffic. Benchmarks measure intelligence per prompt. Your invoice measures intelligence per dollar across an entire workflow, including the failures.
There is also the framing problem. "Near-Astra intelligence" is OpenAI's characterization of its own model, tested against its own eval set, announced on its own blog. Nobody outside OpenAI has independently reproduced that comparison yet. Treat the claim as a starting hypothesis, not a verified fact, until third-party benchmarks or your own eval harness confirms it on tasks that look like what you actually run.
What "near-flagship at a fraction of the cost" actually requires under the hood
Think of a busy restaurant kitchen. The head chef, Astra, can handle any dish on the menu, including the ones that require improvising when an ingredient is missing. Training and running that chef on every ticket is expensive: more headcount, more attention on each plate, higher prices at the register. Sol is the kitchen's answer to that cost problem: a smaller team of line cooks who have been trained specifically on the dishes that make up 90 percent of the orders, with a much lighter footprint per plate. They are fast and cheap because they are not paying the overhead of being ready for everything.
In practice this usually means a smaller active parameter count, more selective routing inside a mixture-of-experts architecture, or a distillation process where the flagship model's outputs are used to train a leaner one to mimic its behavior on common request patterns. The result looks nearly identical on the requests it was tuned for. It looks noticeably different the moment you send it something actually novel, like a reasoning chain the training data never resembled.
For anyone building against the API, the practical test is trivial to run. Swap the model string and diff the outputs on your actual traffic:
curl https://api.openai.com/v1/chat/completions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6.1-sol",
"messages": [{"role": "user", "content": "your real prompt here"}]
}'
Run the same prompt against gpt-6-astra and compare. If you cannot tell the difference on your workload, you have your answer. If you can, you have found exactly the boundary where "near" stops applying.
Sol, Astra, and the closest priced competitor
The most useful comparison is not Sol against its own sibling. It is Sol against whatever a team would otherwise reach for at a similar price point, which right now is Claude Sonnet 5.5.
| Model | Input / 1M tokens | Output / 1M tokens | Released | Best fit |
|---|---|---|---|---|
| GPT-6.1 Sol | $2.00 | $10.00 | 2026-09 | High-volume production traffic where task complexity is mostly predictable |
| GPT-6 Astra | $10.00 | $50.00 | 2026-09 | Novel, high-stakes, or open-ended reasoning where a wrong answer is expensive |
| Claude Sonnet 5.5 | $2.00 | $10.00 | 2026-09 | Teams already inside the Anthropic ecosystem who want a like-for-like price comparison |
For a startup shipping customer-facing chat at volume, Sol at parity pricing with Sonnet 5.5 is the sensible default and worth an eval run before committing. For anyone running agentic pipelines with long tool-calling chains where a single bad reasoning step compounds, Astra's higher ceiling is still the safer bet even at 5x the cost. For teams already committed to Claude, this announcement changes nothing except giving you a cleaner apples-to-apples number to cite when your manager asks why you have not switched providers.
How to actually test the "fifth of the price" claim before you migrate
- Pull 50 to 100 real production prompts from your logs, not synthetic test cases. Include your hardest 10 percent, not just the median request.
- Run each prompt against both
gpt-6-astraandgpt-6.1-sol, capturing the full response and the token count for both. - Score the outputs against your existing acceptance criteria, whatever you already use to judge whether a response shipped or got rejected.
- Calculate cost per accepted response, not cost per token: total spend divided by the number of outputs that actually passed your bar.
- If Sol's cost-per-accepted-response is meaningfully lower and its rejection rate is within a few points of Astra's, migrate the traffic that matches your test set. If the rejection rate gap is wide, keep Astra for that segment and route only the easy cases to Sol.
Verification test: take the 10 hardest prompts from your sample, the ones closest to the edge of what your product does today, and confirm Sol's acceptance rate on just that subset before rolling out anything broader. That subset is where "near-Astra" either holds up or does not.
Why Multi-Step Tool-Calling Loops Erase the Savings
The documented failure pattern with cheaper flagship-adjacent models is not that they produce obviously bad output. It is that they degrade quietly in multi-step agentic contexts: a tool-calling loop that would resolve in three steps on the flagship model stretches to six or seven on the cheaper tier, each step burning tokens on both the mistake and the correction. Teams running ChatGPT Work or custom agent stacks on top of the API have hit this exact issue before with earlier "mini" and "flash" tiers: the per-token savings looked great in a spreadsheet and evaporated once retry loops were counted. If your workflow involves an agent calling tools repeatedly to complete a task, benchmark the full loop, not the first response, before you assume the 5x price cut survives contact with your actual pipeline. For background on how these benchmark-versus-production gaps tend to play out, see our look at separating LLM hype from reality.
Prediction, six months out
By 2027-03-30, expect at least one widely cited third-party eval or agent-framework benchmark to show GPT-6.1 Sol underperforming GPT-6 Astra specifically on multi-step tool-calling tasks by a double-digit margin in success rate, even while matching or beating it on single-turn benchmarks. If no such gap shows up in independent testing by then, this piece's skepticism about "near-Astra intelligence" was wrong.
Tools mentioned in this article
Some links in this article are affiliate links. Learn more.