26 July 2026·6 min readLLM pricingClaude Opus 5

Opus 5 arrived at half Fable's price. How to buy AI so the next price drop lands in your budget

Claude Opus 5 shipped on 24 July at the same $5 and $25 per million tokens as Opus 4.8, and Anthropic says it gets close to Fable 5 at half the cost per task. The price of intelligence keeps falling, but increasingly it falls through fewer tokens rather than a lower list price, and most AI contracts are written so the buyer never sees it.

A
Ajay Dhillon
Founder

Update, 31 July 2026: OpenAI cut API prices for GPT-5.6 Luna by about 80 per cent and Terra by about 20 per cent on 30 July. Two frontier vendors moved the price of capability in the same week, in two different ways. The advice below applies to both.

Claude Opus 5 shipped on 24 July at $5 per million input tokens and $25 per million output. That is exactly what Opus 4.8 cost. Anthropic's claim is that it gets close to Fable 5, its top model, at half the price, and the more interesting claim sits underneath: on the coding and business-task benchmarks it published, Opus 5 does the same work with far fewer tokens. Early-access customers reported the same thing in different words. One quoted 26 per cent fewer tokens at similar quality; another, a third fewer turns and tool calls; a trading desk said its benchmark ran on roughly a seventh of the reasoning tokens Opus 4.8 needed.

That is a price cut. It is just not the kind procurement teams can see, because the rate card did not change. The cost of a completed task fell while the cost of a token held still, and a contract that commits spend in tokens captures none of it.

Three ways the price of intelligence falls

The first is the list price. Sonnet 5 launched at $2 and $10 in June against Sonnet 4.6's $3 and $15. OpenAI's cuts on 30 July are the same thing. These are visible, they make headlines, and they are the smallest part of the story.

The second is efficiency at the same price. Opus 5 for Opus 4.8's money, with fewer tokens per task. Effort settings that let you pay for less thinking on easy steps. Cache reads at a tenth of the input rate. Batch endpoints at half price for anything that can wait. None of these change the rate card, and together they matter more than the rate card.

The third is the tier shift. Work that needed the flagship in January runs on the mid-tier in July. Anthropic's own charts show Opus 5 beating Fable 5's best computer-use result at about a third of the cost, and Sonnet 5 reaching Opus 4.8 on some tasks at high effort. The cheapest way to run a workload in 2026 is usually a model that did not exist when the workload was designed.

A buyer who tracks only the first of these will conclude that AI prices are roughly flat. A buyer who tracks all three will find that the cost of a given task has fallen by more than half in a year, and will structure their spending to keep catching it.

What a contract should say

Most enterprise AI agreements were drafted by people who had bought cloud before, and they imported the cloud habit of committing to a volume in exchange for a discount. Committed spend in tokens on a named model is the worst possible shape for 2026, because it fixes the two things most likely to get cheaper.

Commit to spend in dollars, across models, if you commit at all. Then a tier shift or a token-efficient release reduces what you use rather than what you owe.

Ask for a clause that passes list-price reductions through during the term. Vendors will resist, and a shorter term with no commitment is often the better answer.

Keep the right to change models without renegotiating. The contract is with the platform, and the model is a parameter.

And do not pay a premium for Fast mode across the board. Opus 5's Fast mode runs at about two and a half times the speed for twice the price, which is a bargain for a customer-facing step with a latency budget and a waste for a batch job overnight.

What the architecture should allow

A contract can only capture a price drop the system is able to use. Four properties decide that.

Per-step model selection. Every step in an agent's loop names its own model and effort level, and both are configuration rather than code. When a cheaper model passes the tests for a step, the change is a line in a file.

An evaluation set per workload. The reason most teams stay on an expensive model is that they cannot prove the cheaper one is safe. Two hundred real cases with a grader turns "we think it is fine" into a number, and the number is what lets you move within days of a release instead of quarters. We have written about building that set at length because it is the single asset that makes every other saving reachable.

Caching by design. Static context, such as policy documents, schemas and tool definitions, placed at the front of the prompt where the cache can hold it. Anthropic's new mid-conversation tool changes, released alongside Opus 5, let an agent swap tools without invalidating that cache, which removes one of the common reasons caching failed to deliver.

Cost per task in the logs. If you cannot see what a task cost last week and this week, you cannot tell whether a release helped. Put the number next to the success rate and review both monthly.

Re-baseline quarterly

Set a calendar reminder for every quarter that says: for each production workload, what is the cheapest model and effort setting that passes our evaluation set today, and what would it save. In the last year that question has had a different answer every quarter for nearly every workload we run, and the answer has rarely been the model the workload started on.

The vendors are competing on the cost of a completed task now, whether or not they say so on the pricing page. Buying AI well in 2026 means being the kind of customer who can take the deal when it arrives, which is a matter of contract shape and system shape rather than negotiating skill. We work through both in the first fortnight of an AI strategy engagement, and the contract review is usually the shorter half.

Frequently asked

How much does Claude Opus 5 cost? $5 per million input tokens and $25 per million output tokens, the same as Opus 4.8, with Fast mode at twice that price for roughly 2.5 times the speed. Anthropic says Opus 5 gets close to Claude Fable 5's performance at about half the cost per task, mainly by completing work with fewer tokens and fewer tool calls.

Are AI model prices falling in 2026? Yes, in three ways: lower list prices (Sonnet 5 in June, OpenAI's GPT-5.6 Luna and Terra cuts on 30 July), more work per token at the same price (Opus 5 against Opus 4.8), and workloads moving from flagship to mid-tier models as the mid-tier catches up. The cost of a completed task has fallen faster than any rate card shows.

How should an enterprise AI contract be structured? Commit in dollars across models rather than in tokens on a named model, ask for list-price reductions to pass through during the term, keep the right to change models without renegotiating, and pay for fast or priority modes only on the steps that need low latency. Pair the contract with per-step model selection and an evaluation set so the system can actually adopt a cheaper model when one appears.

Related reading

The price of a token is the least informative number on the invoice. Buy against the price of a task, and build so that number can fall.

Written by
Ajay Dhillon · Founder
08 · Start here

Let’sbuildyoursystemnext.

Thirty minutes with someone who’d be doing the work. No slide deck, no intake form. We’ll tell you what’s feasible, where you’ll hit friction, and what we’d pick up first.

Response
< 24 hours
First read
No NDA needed
Bangalore / Remote
UTC ±12