OpenAI shipped a new flagship model on July 9. If you run a mid-sized company, here is the correct amount of your week to spend reacting to that: none.
That's not cynicism about the model. It's what the launch itself says, read closely.
What actually shipped
GPT-5.6 went generally available across ChatGPT, Codex, and the API. It comes in three tiers, and OpenAI's own announcement lists the prices per million tokens: Sol at $5 in / $30 out, Terra at $2.50 / $15, Luna at $1 / $6.
Look at the flagship number for a second. $5 / $30 is exactly what GPT-5.5 cost. The new frontier model launched at the old frontier model's price. That is not what a decisive capability leap looks like when it's priced by the company that built it.
On performance, OpenAI points at the Artificial Analysis Intelligence Index, where Sol lands within a single point of Anthropic's Claude Fable 5 — 59 against 60 — while, in OpenAI's framing, "completing tasks in 61% less time at roughly half the estimated cost." (That's OpenAI's chosen benchmark and OpenAI's chosen comparison, so read the speed and cost claims as a vendor's best case. The part worth keeping is the gap: one point.)
One point. Between the two most expensive frontier models in the industry, from the two labs furthest ahead. That's the whole story, and everything below is a footnote to it.
The tell: OpenAI didn't switch its own default
Here's the detail that got almost no coverage, and it's the one that should decide how you feel about all of this.
GPT-5.6 did not become the default model in ChatGPT. GPT-5.5 Instant is still what everyday conversations run on. Sol powers the reasoning options on eligible paid plans — an option you select, not the thing you get.
Sit with that. The company that built the model, that has every commercial reason to put its newest and most impressive work in front of every user, looked at its own launch and decided the previous model was still the right default for most requests.
If OpenAI isn't in a hurry to move everyone onto GPT-5.6, the case for you rebuilding anything around it this quarter is thin.
What companies actually do with the "best" model
The buyers furthest along have already stopped shopping this way.
Forbes reported in July on why leaderboards no longer decide enterprise AI buying, and the numbers in it are more useful than any benchmark chart. Databricks ran a test on real million-line codebases — not puzzles, actual work. Claude Opus 4.8 completed 87% of tasks at $1.94 each. An open-weight model, GLM-5.2, came in statistically tied at $1.28.
Same outcome. A third less money. From a model most people reading this couldn't name.
That's not an isolated result, either. Per the same reporting, Chinese-origin models climbed from 4.5% to somewhere between 30% and 46% of enterprise token volume on OpenRouter. And Microsoft — which has spent more on a relationship with OpenAI than most companies are worth — now routes commodity tasks to its own in-house models and saves the frontier models for the genuinely hard reasoning.
Read that last one twice. Microsoft's answer to "which model should we use?" is "depends on the task, and mostly not the expensive one." They're not picking a winner. They're routing.
Commodity doesn't mean simple
This is where the "models are commoditizing" argument usually overreaches, so let's not.
Cheap and interchangeable does not mean easy. If anything it's the opposite: the more real options exist, the more someone has to own the decision of what runs where, and when, and what happens when it fails.
DeepSeek is the sharpest example. Its V4 release was reported in mid-July to introduce something genuinely new: time-of-day surge pricing, with rates roughly doubling during Beijing peak hours. (Reported, but we couldn't confirm it against DeepSeek's own documentation before publishing this — treat it as directional until you verify it yourself.) Even as a directional signal, think about what it implies. The cheapest frontier-class inference on the market may now cost different amounts depending on what time it is in another hemisphere.
Google, meanwhile, is a lesson in a different direction: Gemini 3.5 Pro is still in limited preview after missing multiple general-availability targets. Any plan built around a model that was promised but hasn't shipped is a plan built on someone else's roadmap.
So: the models converged, the prices fell, the options multiplied — and the job got harder, not easier. That job has a name. It's not model selection. It's deployment.
Update — July 21, 2026: the loose ends resolved, all in the same direction
In the days after this published, the two things we told you to watch both settled — and a third piece of evidence landed. All three point the same way.
The DeepSeek "surge pricing" never showed up. Above, we flagged the reported time-of-day surcharge as unconfirmed and told you to check it against the source. We did. DeepSeek's own API pricing page lists flat rates — deepseek-v4-flash at $0.14 in / $0.28 out, v4-pro at $0.435 / $0.87 — and describes the cost as nothing more exotic than "number of tokens × price." There is no peak-hour surcharge anywhere on it. The scary pricing mechanic that made the rounds simply isn't real at the source. Which is the smaller lesson inside the bigger one: the noise about any single model's pricing rarely survives contact with the vendor's actual docs.
Google still hasn't shipped Gemini 3.5 Pro. We noted it was overdue; it's now more overdue. A mid-July general-availability target came and went with no model card, no API listing, and no published pricing — Google's public model list still shows only Gemini 3.5 Flash as generally available. (As recorded in July — three further Flash generations have shipped since; see the September 2026 update.) A frontier model that keeps not arriving is the cleanest argument we could ask for against welding your operations to one vendor's roadmap. You can't deploy a promise.
And the "route by cost, not leaderboard" behavior now has a name attached at the very top. In a July 9 analysis, industry analyst Bertrand Duperrin quoted Coinbase CEO Brian Armstrong describing the exact playbook we laid out — the company is "automatically routing requests to the least expensive model compatible with the task at hand, caching responses, and limiting calls to the most expensive models when they do not provide tangible benefits." Duperrin's summary of where value is going could be the subtitle of this post: it is "shifting from the models themselves toward their integration into systems, data governance, agent orchestration."
Two weeks, three data points, one direction. Nothing here changes the play. It just keeps proving it.
Update — August 3, 2026: the prices moved a lot, the conclusion didn't — and we owe you a correction
Thirteen days on, the price table at the top of this post is wrong in four places. That isn't an embarrassment; it's the argument. Here's the current state — starting with the thing we got wrong.
The correction: DeepSeek's surge pricing is on the page now
In the July 21 update above we wrote that DeepSeek's pricing page listed flat rates and that "there is no peak-hour surcharge anywhere on it." We checked again on August 3. Now there is. DeepSeek's pricing page currently states: "The DeepSeek API service will soon adopt a peak/off-peak pricing policy. During peak hours, prices will be 2x the regular prices, applicable to all billing items." Peak is given as 9:00–12:00 and 14:00–18:00, Beijing time.
Read the tense before you react, because it still isn't what the original reporting claimed. It says "will soon adopt," not "has adopted." The listed rates are still flat — v4-flash at $0.14 in / $0.28 out, v4-pro at $0.435 / $0.87 — and on timing the page offers only that the effective date "will be subject to the official announcement." So the mechanic is real and announced, it is not in force, and no one outside DeepSeek knows when it starts. What we published in July was right about the rates and too absolute about the page. We're correcting it down here rather than quietly editing the paragraph above, so you can see what moved.
OpenAI cut its own prices — and the flagship still didn't move
On July 30, OpenAI repriced the GPT-5.6 line. Measured against the launch numbers quoted at the top of this post, the current rates are Luna at $0.20 in / $1.20 out (was $1 / $6 — an 80% cut), Terra at $2.00 / $12.00 (was $2.50 / $15), and Sol unchanged at $5 / $30. A new Fast mode runs at roughly double the standard rate for about 2.5× the speed.
Sit with the shape of that. Six weeks after launch, the budget tier costs a fifth of what it did, the middle tier is down a fifth, and the flagship hasn't moved a cent. The expensive model is the one holding its price; everything underneath it is falling out from under it. That is what commoditization looks like from the inside, and it is moving faster than any procurement cycle you could run against it.
Anthropic shipped a better flagship at the same price — and put a clock on the cheap one
Claude Opus 5 became generally available on July 24 at $5 in / $25 out — identical to Opus 4.8. More capability, same price, which is the cheapest kind of upgrade to get signed off: nobody has to reopen a budget line.
The one to put in the diary: Claude Sonnet 5's $2 / $10 is introductory pricing that runs only through August 31, 2026, after which the list rate is $3 / $15. Anyone modelling next year's run-rate on today's Sonnet number is understating it by half. And Claude Opus 4.1 retires on August 5 — if anything you own has that model ID hard-coded, it stops working this week. Which is a fairly pointed illustration of question 4 below.
Google shipped the Flash tier and still hasn't shipped the Pro one
Gemini 3.6 Flash went generally available on July 21 at $1.50 / $7.50 — an output-price cut from 3.5 Flash's $9.00.
Gemini 3.5 Pro is still not generally available, and we can now attach Google's own words to that rather than inferring it from an empty pricing page: their July 21 announcement describes it as "currently testing with partners." It remains absent from the Gemini API model list and from the published pricing. We flagged it as overdue on July 16, again on July 21, and here we are on August 3.
Three consecutive updates in which the model that would supposedly settle the "which is best" question has not arrived. If you deferred a decision waiting for it, you have now been waiting since May.
What changed in the conclusion: nothing
Four price changes, a new flagship from each of two labs, one deprecation with a hard date, one pricing mechanic announced without one, and one model that still hasn't shipped — in thirteen days. If you had picked a model on July 21 and welded your systems to it, you would already be either repricing or migrating.
That's the whole case for treating the model as a swappable component. It isn't a philosophical position about AI. It's an observation about how often the table under your feet gets rewritten.
Update — August 2026: a scheduled price rise was cancelled, and the other side moved too
Two changes worth knowing about if you are mid-evaluation, both read off the vendors’ own pricing pages rather than coverage of them.
Anthropic cancelled a price rise it had already announced. Sonnet 5 launched at $2 per million input tokens and $10 per million output, described at the time as introductory pricing through 31 August 2026, with a rise to $3/$15 scheduled for 1 September. That rise is off. The pricing page now says, in its own words, that the $2/$10 pricing “is now the standard price” and that “the previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur.”
Note the wording, because it is doing work: standard price, not permanent. A standard price is one a vendor can change with notice. What has actually happened is that a dated increase was withdrawn, which is better than it sounds and less than a promise.
On the other side, OpenAI’s pricing page currently lists gpt-5.6-sol at $4.00 per million input and $20.00 per million output — and this is the part most write-ups drop — that is the short-context tier. The long-context tier is $8.00 and $30.00. Same model, double the input price, depending on how much you put in the window. If you are comparing a headline number against Anthropic’s, make sure you are comparing the tier you will actually be running in, because a long-context RAG workload prices very differently from a short prompt.
⚠️ We are deliberately quoting the levels and not narrating a price cut. Reuters reported a reduction taking effect in late August; OpenAI’s own page carries no dated change note, and we do not publish a price movement we cannot see at the source. The numbers above are what the vendor lists today. Check them again before you commit to anything — that is the whole lesson of the August 3 correction above.
Does any of this change the recommendation? No, and that is the point. The gap between these two is now small enough that it should not be what decides your architecture. What decides it is which model your workload actually needs, whether you can swap it later without rewriting everything, and whether the thing calling it is built so that a price move is a config change rather than a project.
Update — September 2026: two flagships, one price, and a change no rate card shows
Three frontier launches inside three days. OpenAI shipped GPT-6 Astra alongside a repriced GPT-5.6 family. Anthropic shipped Claude Fable 5.1. Google shipped Gemini 3.8 Flash. Start with the thing that is hard to unsee: OpenAI’s new flagship and Anthropic’s new flagship both list $10 in / $50 out per million tokens. Two vendors, two independent launches days apart, the same headline number. We are not going to speculate about why. We are just going to note that the number everyone was waiting to argue about turned out to be the same number on both sides.
The current OpenAI Standard tier list, per million tokens, in / cached / out, short context then long context:
- gpt-6-astra — $10.00 / $1.00 / $50.00; long context $20.00 / $2.00 / $75.00
- gpt-5.6-sol — $4.00 / $0.40 / $20.00; long context $8.00 / $0.80 / $30.00
- gpt-5.6-terra — $2.00 / $0.20 / $12.00; long context $4.00 / $0.40 / $18.00
- gpt-5.6-luna — $0.20 / $0.02 / $1.20; long context $0.40 / $0.04 / $1.80
OpenAI’s model page says Astra “is rolling out today for enterprises in our Trusted Access Program, with access through API and our Plus, Pro, Business and Enterprise plans coming in the coming days.” Context window 1,050,000 tokens, knowledge cutoff April 30, 2026. Its supported-endpoints list excludes Assistants. If your integration runs on that endpoint, the flagship you just read about is not available to you today, whatever the price says. That is the deployment surface deciding what you can use, not the benchmark.
On the Anthropic side, per million tokens: Claude Fable 5.1 at $10 base input and $50 output, with cache hits at $0.25. Claude Mythos 5.1 lists the same and is marked limited availability. Claude Fable 5 lists $10 / $50 with cache hits at $1. Claude Opus 5 and Opus 4.8 list $5 / $25. Claude Sonnet 5 lists $2 / $10. A footnote explains the cache gap: “Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. All other models use the standard 0.1x multiplier.”
Note what that comparison is and is not. OpenAI’s $10 / $50 is its short-context rate; Anthropic’s $10 / $50 is its only rate. Which brings us to the part that matters more than any launch price.
The first real difference is long context, and the two vendors have priced it in opposite directions. OpenAI states that “prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.” Anthropic states that “Claude 4.6 and later models … include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.)” If your workload is long documents, that structural choice will decide your bill long after the headline rate stops being interesting.
The second is better, and it is the most useful thing in this update. Anthropic’s pricing page states: “Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape.” Anthropic frames that as a property of the tokenizer, not as a price change, and we are not going to reframe it. But read what it means operationally: the rate card can sit still while the same document turns into more billable units. We are not going to put a percentage on your bill either, because the page says the increase depends on the content and the workload. That is the whole argument of this post in one footnote. The number everyone compares is not the number that decides what you pay.
Google, meanwhile, keeps doing the thing this post has been pointing at since July. Gemini 3.8 Flash lists introductory pricing of $0.75 in / $3.75 out through December 31, 2026, and standard pricing of $1.50 / $7.50 from January 1, 2027. That is the third Flash generation in roughly six weeks — 3.6, then 3.7, then 3.8. And “Gemini 3.5 Pro” still does not appear on that pricing page at all: not preview, not generally available, not listed. Google has shipped Flash generations instead. We said in August that a frontier model which keeps not arriving is an argument against welding your operations to any one vendor’s roadmap. Two more Flash generations and still no Pro is not a reason to soften that. It is the case getting stronger.
Everything above this section was re-checked against the same three pricing pages while writing it, and the earlier figures hold: Sol’s split between short and long context is as recorded, and Anthropic’s cancelled increase is still cancelled. Nothing in eight weeks of price moves has changed what we would tell you to do. Which, four updates in, is starting to look less like a coincidence and more like the finding.
Two days later, one sentence above is out of date, and the way it expired is the point. We quoted OpenAI saying Astra was rolling out to Trusted Access enterprises with broader access “coming in the coming days.” Those days have passed. Microsoft’s own blog now says Astra “is now generally available for all customers in Microsoft Foundry”, and GitHub’s changelog dated September 4 says “GPT-6 Astra is generally available in GitHub Copilot”, available to “Copilot Pro+, Max, Business, and Enterprise users.” OpenAI’s own pages remain unreadable to us, so we are quoting the two vendors who shipped it rather than the one who built it.
Now the part worth stopping on, because it is not a pricing question and it is not a benchmark question. Here is how Astra arrived in Copilot, in GitHub’s words:
“new models are enabled automatically unless an administrator has turned off the global default or explicitly disables this model.”
Read that as an operator rather than as a reader. A new frontier model appeared in your developers’ tooling this week, billed at provider list pricing under usage-based billing, and the only thing that would have stopped it is a decision somebody made before the model existed. Not a decision about Astra — a standing decision about defaults, taken in advance, by someone who knew this was how the pipe works.
That inverts the question this whole post is built around. We have spent four updates arguing that picking a model is the wrong unit of decision because the answer changes every few weeks. This is the sharper version: for the tools your team already has, you may not be picking at all. The vendor picks, on a schedule you do not control, and your position defaults to whatever was set the last time anyone looked.
None of that is a complaint about GitHub, and the setting is documented and reachable — Copilot Business and Enterprise administrators manage it through the model policy in Copilot settings. It is a comment on where the decision actually lives. If your answer to “which model are we using?” is the name of a model, you are answering a question nobody asked. The answer that survives contact with a launch week is a policy: who decides, how fast changes land, and what happens by default when nobody is watching.
The question that actually matters
"Which AI is best?" is a question with a fresh answer every six weeks and no operational consequence. Here's what to ask instead:
- What work are we actually trying to move? Not "where can we use AI" — which process eats the most hours for the least judgment. The inbox. The phone after 5pm. The report someone rebuilds every Monday.
- Is it wired into our tools and our data? A model that can't see your CRM, your calendar, or your ticket history is a very articulate stranger. The integration is the product.
- What happens when it's wrong? Who reviews it, what does it escalate, where does a human step in. If there's no answer, you don't have a system — you have a demo.
- Can we swap the engine without rebuilding the car? This is the real one. Given how fast the frontier moves, any system that's welded to one vendor's model is a system you'll rebuild inside a year. The same argument applies one layer down, to the plumbing: the standard that connects agents to your tools just had a breaking rewrite, with a twelve-month clock on the deprecated parts.
Notice none of those mention a model name. That's the point. The model is the easiest, cheapest, most replaceable component in the entire stack — and it's the only one anybody argues about.
We build on the frontier models, all of them, and we treat that choice as an implementation detail we own and can change. When something better ships, we move — you shouldn't have to know, or care, or re-sign anything. What you should notice is that the work keeps getting done. That matters most on the days a provider degrades rather than improves — seven major AI incidents in nine days is what that looks like from outside, and why being able to move a request beats having picked the right model.
Because the model was never the hard part. Getting it plumbed into the way your company actually runs — that's the hard part. That's the part that's worth paying for. And it doesn't have a version number.
Related reading: From ChatGPT Chaos to Integrated AI Systems — why a chat tab saves an individual 30 minutes but never changes how the company operates, and why a third of companies that cut jobs for AI are hiring them back.
Related service: AI integration — choosing what to build is most of the job, and we take no margin on whichever model wins.