6 September 2026·7 min readGPT-6 AstraClaude Fable 5.1

GPT-6 Astra, Claude Fable 5.1 and a $12.9 billion Hugging Face deal in one week. What changed for buyers

Anthropic shipped Fable 5.1 on 1 September, OpenAI shipped GPT-6 Astra on 3 September, and NVIDIA agreed to buy Hugging Face the same day. Both frontier models now list at $10 and $50 per million tokens. Underneath the headlines, four things moved for anyone buying AI: where the real price cuts are, how the top models now differ, who owns the open-model supply chain, and what the new safety tiers mean for access.

A
Ajay Dhillon
Founder

Three announcements in three days.

On 1 September Anthropic released Claude Fable 5.1 and Claude Mythos 5.1, the gated version for vetted cybersecurity and life-sciences work. The per-token price held at $10 per million input and $50 per million output, and the price of a cache read fell by 75 per cent to $0.25, which Anthropic estimates cuts typical workloads by about a quarter and heavily agentic ones by up to 45 per cent.

On 3 September OpenAI released GPT-6 Astra, at $10 and $50 per million tokens, calling it the most intelligent and aligned model it has built. It is the first OpenAI model to reach the Critical level of cybersecurity capability under the company's Preparedness Framework, it was delayed to add safeguards after July's Hugging Face incident, and it rolled out first to a limited group of organisations before wider availability.

The same day NVIDIA agreed to acquire Hugging Face for $12.93 billion, including a retention programme for staff, and Jensen Huang said the platform would stay open and would not require NVIDIA hardware.

Each announcement got its own week of coverage. Read together, they settle four questions a buyer has been asking all year.

One: the list price is no longer where the cuts happen

The two frontier models now cost the same per token, and neither company cut that number. Anthropic cut the price of the cache read. OpenAI's model, according to independent measurement by Artificial Analysis, uses about a third as many tokens as GPT-5.6 Sol on coding-agent evaluations, which brings its cost per coding task roughly level with Sol despite a per-token price two and a half times higher. On general one-shot tasks, the same measurement puts Astra about 75 per cent dearer per task than Sol.

This is the pattern we described in July after Opus 5: the price of intelligence falls through efficiency and through caching, and a buyer who tracks only the rate card sees none of it. Fable 5.1's cache cut is close to a straight price reduction for any agent with a stable preamble. Astra's efficiency is a price cut on long agentic runs and a price rise on everything else.

The buying consequence is unchanged and now unavoidable: measure cost per completed task on your own work, per model, per step. A per-token comparison of Astra and Fable 5.1 tells you nothing. They cost the same per token and very different amounts per task, in opposite directions depending on the task.

Two: the frontier has split by capability profile

For most of the last two years the question was which model was best. This week made it clear the question is best at what.

Astra's largest gains over its predecessor are in computer use, long-context retrieval, terminal-based agentic work and cybersecurity, where OpenAI's and independent numbers agree it leads. On a needle-in-a-haystack test across a million tokens it reportedly reached 96 per cent against 74 per cent for Sol. On broad intelligence measures, Artificial Analysis scored it at 61, level with Sol and behind Fable 5.1 at 66, and on its coding-agent index it trails Fable 5.1 as well. OpenAI's headline ARC-AGI-3 score was produced in a provider-specific harness; the neutral harness produced a much lower number.

None of this makes either model wrong. It makes "which model" the wrong unit. An agent that reads a long contract, plans, operates a browser and writes a summary may want Fable 5.1 for the plan and Astra for the browser. Per-step model selection stopped being an optimisation and became the design.

It also means vendor benchmarks are less useful than they were. When two labs each lead on a different set of tests, the only number that resolves the question is your own evaluation set. Teams that built one this year can answer "should we use Astra" in a day. Teams that did not are reading press releases.

Three: the open-model supply chain has an owner

Hugging Face is where open models are published, downloaded and built on. Our August piece drew on its own data: 151,000 Qwen derivatives, the Chinese labs' frontier-scale releases, the licence drift on the largest new models. From now on it is owned by the company that sells most of the hardware those models run on.

Huang's commitments are specific: the platform stays open, developers choose their models, frameworks, clouds and hardware, and NVIDIA hardware will not be a requirement. Forrester's read is that NVIDIA gains control and visibility over a key layer of the AI supply chain, and that this could improve provenance, vulnerability disclosure and guardrails at scale, while concentrating influence over a critical piece of the open ecosystem. Both readings are true.

For a buyer who self-hosts, two practical steps. Keep local copies of the weights you depend on, under the licence you downloaded them with, dated. And treat the Hub the way you treat any critical vendor: a dependency with an owner, a roadmap you do not control, and terms that can change. The acquisition needs regulatory approval, and the commitments are made by the buyer rather than written into the platform.

Four: capability tiers now come with access tiers

Astra is the first OpenAI model at the Critical cyber threshold. Anthropic's Mythos line exists precisely because Fable's capabilities in cyber and biology are gated behind verification. OpenAI has added misalignment monitoring to all tool-using inference for Astra, at what it calls significant compute cost, and says the model is more capable of controlling its own reasoning traces and less likely to include incriminating detail in them, which reduces how well those traces can be monitored.

The pattern for buyers is that the top of the market is moving to tiered access: full capability behind identity verification and usage agreements; broad availability with safeguards that will sometimes refuse legitimate work (though Anthropic says Fable 5.1 produces fewer false positives than Fable 5, and now routes blocked biology requests to Opus 5 rather than failing them); and enterprise agreements that include monitoring the vendor runs on your traffic.

Plan for it. If your work touches security research, life sciences or anything a classifier might mistake for either, apply for the verified programmes now rather than when a project is blocked. Ask vendors what monitoring runs on your inference and what it retains. And keep a fallback model in the loop for the requests a safeguard declines, which Anthropic now supports automatically on the API.

What to do this month

Add Astra and Fable 5.1 to your evaluation runs, per step, and record cost per task alongside accuracy. Move any static preamble into cache position if it is not already there; the Fable 5.1 cut makes this the single largest saving available on that model. Snapshot the open weights you depend on. And if you run agents that operate browsers or terminals, test Astra on that step specifically, because that is where its lead is.

One more signal from the same week is worth keeping. Reports surfaced that ServiceNow had begun monitoring employees' usage after consuming its annual Anthropic budget faster than planned. A company of that scale finding its AI budget gone mid-year is the clearest evidence yet that cost governance, per task and per agent, is now a board-level control rather than an engineering nicety. It is the first thing we look at in an AI strategy engagement, and this week made the case for us.

Frequently asked

How much do GPT-6 Astra and Claude Fable 5.1 cost? Both are priced at $10 per million input tokens and $50 per million output tokens. Anthropic cut Fable 5.1's cache-read price by 75 per cent to $0.25 per million, which it estimates reduces typical workload costs by about 25 per cent and agentic workloads by up to 45 per cent. GPT-6 Astra is about 2.5 times the per-token price of GPT-5.6 Sol, but independent measurements found it uses roughly a third as many tokens on coding-agent tasks.

What is GPT-6 Astra best at? Independent and vendor benchmarks agree its largest gains are in computer use, long-context retrieval, terminal-based agentic workflows and cybersecurity. On broad intelligence and coding-agent indices from Artificial Analysis it scores level with GPT-5.6 Sol and behind Claude Fable 5.1, so the right choice depends on the specific step in a workflow.

What does NVIDIA buying Hugging Face mean for companies using open models? NVIDIA agreed on 3 September 2026 to acquire Hugging Face for $12.93 billion, subject to regulatory approval, and has said the platform will remain open with no requirement to use NVIDIA hardware. Companies that self-host open models should keep dated local copies of the weights they depend on, record the licence they downloaded under, and treat the Hub as a critical vendor dependency.

Related reading

Two frontier models at one price, an open-model platform with a new owner, and a capability tier that now comes with a verification form. The week changed less about which model is best than about how a company has to buy, and the companies with their own evaluation sets and cost-per-task numbers were the only ones who could act on it by Friday.

Written by
Ajay Dhillon · Founder
08 · Start here

Let’sbuildyoursystemnext.

Thirty minutes with someone who’d be doing the work. No slide deck, no intake form. We’ll tell you what’s feasible, where you’ll hit friction, and what we’d pick up first.

Response
< 24 hours
First read
No NDA needed
Bangalore / Remote
UTC ±12