12 July 2026·6 min readBuild vs buyChatGPT Work

ChatGPT Work, Copilot, or your own stack: where enterprise AI money should go

OpenAI launched ChatGPT Work and the GPT-5.6 family on 9 July, and the pricing tells you where the industry is heading: a seat for the assistant, metered credits for the agent. That settles most of the buy-versus-build argument. Here is how to split the budget three ways, and the four tests for when building is the right call.

A
Ajay Dhillon
Founder

Update, 31 July 2026: On 30 July OpenAI cut API pricing for GPT-5.6 Luna by about 80 per cent and Terra by about 20 per cent. The launch prices quoted below are now higher than what you would pay. The argument does not change; the build threshold moves a little further out.

On 9 July OpenAI took GPT-5.6 to general availability in three sizes, Sol, Terra and Luna, priced at launch at $5, $2.50 and $1 per million input tokens. The same day it launched ChatGPT Work, an agent that takes on longer tasks across connected apps and files, available on every plan on desktop. The old app was renamed ChatGPT Classic.

The interesting part is how it is billed. ChatGPT Business stays at $20 a seat a month on an annual plan. ChatGPT Work, Codex and the Workspace Agents draw on a shared pool of credits measured in tokens, on top of the seat. OpenAI's own guidance is that Codex runs at roughly $100 to $200 per developer per month depending on usage. Microsoft has the same shape with Copilot seats plus Agent 365 and pay-as-you-go agents. Anthropic tried to split its subscriptions the same way in May and has said a revised version will return.

Every major vendor has landed on the same model. A flat seat for the assistant a person talks to. Metered usage for the agent that works on its own. That structure answers the buy-versus-build question better than most strategy decks do, because it tells you there are three budget lines, and a different logic for each.

Line one: seats, for everyone

The assistant in a chat window is now a commodity, and a cheap one. $20 to $30 a seat buys a frontier model, document upload, connectors to mail and files, and enough governance for most companies. There is no case for building this yourself and there has not been for a year.

The only decisions here are which vendor and how many seats. Which vendor mostly follows where your documents already live: a Microsoft 365 estate leans to Copilot, a Google Workspace estate to Gemini, and ChatGPT and Claude compete on model preference and on the agent products around them. How many seats is the wrong question; the right one is what stops people using the seats they have, which is usually policy uncertainty rather than licence cost.

Budget it like email: everyone gets one, and the conversation ends there.

Line two: credits, for agents on the vendor's platform

This is the new line, and the one finance teams are not used to. ChatGPT Work, Codex, Copilot agents and Claude's programmatic use are all metered, which means the bill scales with what the agents do rather than with how many people you employ.

The right tool here is a budget per workload with an owner, exactly as you would treat cloud spend. A coding agent at $150 a developer a month is a clear decision when set against what a developer costs. A research agent that ran up $4,000 in a month because nobody set a cap is a different conversation. The budgeting method is cost per completed task, measured, with caps enforced in the platform's own controls.

Use this line for anything the vendor's agent can already do well against the systems it already connects to: mail, calendar, documents, spreadsheets, code repositories, the popular SaaS tools. That covers a great deal. It does not cover your ERP from 2009, your regulated data, or the workflow that is the reason customers pay you.

Line three: build, for the two workflows that are yours

Building means calling the models through the API, inside a system your team owns, with your own evaluation set, your own permissions and your own logs. It is more expensive per workflow than a seat and cheaper per task than credits at volume, and it is the only option when the workflow has to be exactly yours.

Four tests, and a workflow should pass at least two before it gets built.

The systems test. The work touches systems the vendor agents cannot reach: an on-premise ERP, a core banking platform, a government registry, a telephony stack. No connector, no vendor agent.

The data test. The data cannot leave the boundary you have agreed with a regulator or a customer. Vendor agents run in the vendor's cloud. Your build runs where you say, including on-premise or in a sovereign cloud.

The volume test. At high volume, direct API pricing per task beats credit rates and you want per-step model selection, caching and batch endpoints, none of which a seat-based product exposes. The invoice pipeline that handles 5,000 documents a month costs under a thousand dollars in model calls when built and considerably more when run through a general-purpose agent's credits.

The product test. The output is what customers buy from you. A law firm's contract review, a lender's underwriting, a distributor's demand forecast. A capability that defines the business should not depend on a vendor's roadmap, and its evaluation set is an asset the company should own.

Most companies have two or three workflows that pass, and almost none have twenty. If the list is longer, it is a list of things that belong on line two.

A worked split

For illustration, a 500-person professional services firm might land here. Seats for all 500 at roughly $25, so about $150,000 a year, which is the smallest line and the least worth arguing about. Credits for perhaps 60 developers on coding agents and 100 knowledge workers on agent products, budgeted per team and reviewed monthly, somewhere between $150,000 and $250,000 a year depending on how hard they are used. And one or two built systems, each costing more in the first year than either of the other lines, and each justified by a specific number: hours removed from a delivery process, days removed from a billing cycle, a service the firm can now sell.

The mistake is to treat this as one decision. Companies that go all-seats never get the agents. Companies that go all-build spend a year rebuilding what a $20 seat provides. The three-line budget is the boring answer and it is the one that works.

What the July launches change

Two things. The credit line got cheaper and more capable at the same time, which pushes more workloads onto line two and raises the bar for building. And the vendors have shown their hand on pricing: the assistant is a loss leader for the agent, and the agent is metered because that is where the compute goes. Plan for the metered line to grow faster than the seat line every year from here.

This is the split we work through in the first week of every AI strategy engagement, and the output is usually a shorter build list than the client walked in with.

Frequently asked

What is ChatGPT Work? ChatGPT Work is an agent OpenAI launched on 9 July 2026 for longer, multi-step tasks that run across a user's connected apps and files. It is available on every ChatGPT plan on desktop, runs on the GPT-5.6 family, and on Business and Enterprise plans is billed from a shared credit pool measured in tokens, on top of the per-seat licence.

Should a company buy Copilot or ChatGPT seats, or build its own AI? Usually all three, on separate budget lines. Seats for every employee for the general assistant. Metered credits, with caps and owners, for vendor agents working against systems they already connect to. A build only for the two or three workflows that touch systems vendors cannot reach, involve data that cannot leave your boundary, run at high volume, or are the product your customers pay for.

How much does an enterprise AI seat cost in 2026? ChatGPT Business is $20 per user per month on an annual plan. Microsoft 365 Copilot and comparable enterprise seats sit in the $20 to $30 range. Enterprise tiers from OpenAI and Anthropic are quoted rather than listed. Agent usage such as ChatGPT Work, Codex and Copilot agents is metered separately, and OpenAI guides that Codex typically runs $100 to $200 per developer per month.

Related reading

Buy the assistant for everyone, meter the agents by workload, and build the two things that make you different. Anything else on the build list is a seat with extra steps.

Written by
Ajay Dhillon · Founder
08 · Start here

Let’sbuildyoursystemnext.

Thirty minutes with someone who’d be doing the work. No slide deck, no intake form. We’ll tell you what’s feasible, where you’ll hit friction, and what we’d pick up first.

Response
< 24 hours
First read
No NDA needed
Bangalore / Remote
UTC ±12