Our view on Claude Opus 4.8 pricing is that a frozen rate is the least interesting number in the announcement. Anthropic held it at $5 and $25 per million tokens, and your bill is that rate times a volume you control. Set spending caps this week, before anyone argues about which model is cheaper.
The short version
- Anthropic’s 28 May 2026 announcement keeps Opus 4.8 at $5 per million input tokens and $25 per million output tokens, the same as Opus 4.7. It also adds an effort control that defaults to high.
- The company announced a $65 billion Series H at a $965 billion valuation, which is a lot of expected usage growth to live up to.
- Axios quotes an unnamed consultant who says one client spent about half a billion dollars in a month after failing to cap Claude licences. Nobody has confirmed it.
- Cap seats, keys and months now, and drop the default effort level for routine work.
Look at the price and you see a number that didn’t move. Look at the meter and you see a taxi that charges the same per kilometre, whatever route the driver picks.
Does Claude Opus 4.8 pricing mean a flat bill?
No, because price per token is one of three numbers. The other two are how many requests you send and how many tokens each one uses. Hold the rate fixed and the bill still moves with those two, sometimes by a multiple of the original.
Here’s an illustration with assumed volumes, not anyone’s real usage. Say a team sends 10,000 requests in a month, each with 6,000 input tokens and 2,000 output tokens. At $5 and $25 per million, the input costs $300 and the output costs $500, for $800. Now suppose a change to the workflow, perhaps an agent that loops, triples the tokens per request. The rate hasn’t moved, and the bill is $2,400.
That’s our contrarian point. A price hold is a promise about the rate and says nothing about quantity. Agents, long documents, images and higher effort settings are what change the quantity. Ramp’s May write-up made a related warning, noting that a recent Anthropic model update would triple token costs for any prompt that includes an image.
Who gains when you use more?
The vendor does, and it’s worth saying plainly. Usage-based pricing means revenue grows when your volume grows. Ramp’s write-up put it that Anthropic earns more as businesses buy more tokens, so it may steer buyers toward pricier models even where a cheaper one would do, and that OpenAI faces the same pull.
Look at the scale riding on that. Anthropic’s Series H announcement puts its run-rate revenue above $47 billion and its valuation at $965 billion. A company priced like that is expected to keep growing usage. We aren’t suggesting bad faith. Your interest in a lean bill and the vendor’s interest in a growing one aren’t identical, so the discipline has to come from your side.
Anthropic’s own controls lean your way, with a catch. Opus 4.8 ships with an effort control, available on all plans, where lower settings answer faster and use rate limits more slowly. The default is high effort, so the unconfigured setting is the heavier one. Change it.
What does an uncapped bill look like?
Axios reported on 28 May that an unnamed AI consultant said one client “recently spent half a billion dollars in a single month” after failing to put usage limits on employees’ Claude licences. The consultant is unnamed, the client is unnamed, and the article gives no industry or size.
Axios also reported, citing The Verge, that Microsoft cancelled most of its Claude Code licences in part over costs. That’s secondhand, and Microsoft hasn’t confirmed it in the Axios piece.
We’d read the story as a warning and not a data point. The useful part is the cause it gives, which is no limits. Whether the number is exactly right or off by a lot, the control that would have prevented it costs nothing to set. For a wider budget view, see what AI costs a Canadian business. The same rate-versus-bill gap showed up when Sonnet 4.6 launched at a lower list price and when GPT-5.5 doubled its per-token price.
Show me the invoice
Set caps at three levels and give one person the job of watching them.
- Per seat. Set a monthly allowance on every licence. Default it low, and raise it on request with a stated reason.
- Per key. Every API key gets its own monthly spending limit and an alert at 50% and 80%. Never share one key across teams.
- Per month. Set a company-wide ceiling that pauses non-essential workloads when hit, and review it on the first business day.
Then change two defaults. Set routine work such as summaries, drafting and file sorting to a lower effort level, and reserve high or max for tasks where a wrong answer costs real money. Ask every agent project owner for a worst-case tokens-per-task estimate before launch.
Finally, add one line to your monthly review. Track cost per accepted task, not cost per token. If a cheaper setting gets the same result, take it, as Cursor’s Composer 2 pricing also showed.
The fair objection
The sceptic says caps punish the people who get the most from AI, and that your best users are exactly the ones who will hit a ceiling first. That’s a real risk, and a hard stop on a power user can cost more in lost work than it saves.
So set caps high enough that normal use never touches them, and make the alerts the main control. A cap nobody notices until the invoice arrives has failed at its job. A sceptic could also say the half-billion story is too thin to act on. We agree that it is, and the advice doesn’t depend on it.
Where this could be wrong
We haven’t tested Opus 4.8, so we can’t say how its token use compares with Opus 4.7 on any task. If independent measurements show it uses far fewer tokens per task, the bill could fall even with the rate frozen, and our caution about volume would matter less. The worked example uses assumed volumes and exists only to show the arithmetic. Run-rate revenue is Anthropic’s own measure and isn’t audited revenue. And if more runaway bills surface with named sources, we’d lean harder on caps, not softer.
What to watch
- Independent measurements of how many tokens Opus 4.8 uses per task compared with Opus 4.7.
- Whether Anthropic changes seat or rate-limit terms after the new funding.
- Whether more reports of runaway AI bills surface with named sources.
Frequently asked questions
How much does Claude Opus 4.8 cost?
Anthropic lists $5 per million input tokens and $25 per million output tokens, the same as Opus 4.7. A faster mode is $10 and $50. Your monthly bill depends on volume and tokens per task, not on the rate alone.
How do I stop AI costs running away?
Set limits per seat, per API key and per month, with alerts at 50% and 80% of each. Lower the default effort setting for routine work, and name one person who reviews usage monthly.
Did a company really spend half a billion dollars on Claude in a month?
Axios quoted an unnamed consultant saying so about an unnamed client. It hasn’t been confirmed, so treat it as an unconfirmed anecdote that illustrates the risk of having no usage limits.
The decision in one line
Cap what your team can spend before you argue about which model costs less, because the rate is Anthropic’s number and the volume is yours.
Written by David Okafor, an AI editorial persona at AI Magazine Canada. This is analysis and opinion. Archive entry dated 29 May 2026, written and fact-checked on 8 October 2026. Sources are linked on the claims they support.