The token economy of an engineering team

Updated September 15, 2026 with current pricing and provider reporting.

For many teams, working with AI agents is already the default way to develop software. Most of the engineering teams I hear from work this way rather than writing code by hand. Some run several in parallel, implementing a change while another agent investigates a problem or reviews code.

We are also creating reusable skills and workflows that make this way of working more manageable. As those workflows improve, it becomes easier to give agents more responsibility. From these conversations, adoption still seems to be growing rapidly, both in how often people use agents and how much work they trust them with.

The focus is often on faster delivery and room to experiment. Token consumption gets less attention, especially when subscriptions keep the bill predictable. Within a plan's limits, a busy day with several agents can cost the same as a quiet one, even though the compute behind it looks very different.

But what would that activity cost if every token were metered? The difference between subscription and API pricing offers a useful way to explore how growing AI usage could affect an engineering team's budget.

Parallel engineering workflows connected to a usage meter, with subscription cards and a budgeting notebook on a desk

What AI usage could cost a team of 30

To put numbers to this, consider a team using Claude with Opus and Codex with GPT-5.6 Sol. Having both subscriptions gives developers another option when one tool reaches its limit. Usage can grow within those allowances without changing the subscription bill, although neither plan guarantees a fixed daily token count.

For an illustrative team of 30, Claude Max 20x and the $200 ChatGPT Pro tier, including Codex, come to $400 per developer per month.

That is $12,000 per month, or $144,000 per year, before taxes and extra usage. This assumes existing individual subscriptions: OpenAI paused new purchases and upgrades to Pro $200 on September 10.

To get a sense of the volumes involved, I asked developers about their average daily token consumption. Those conversations led me to a rough working estimate of 50 million uncached input tokens and 200,000 output tokens per developer per day, across both tools. Usage varied considerably, so this is an informal estimate rather than a measured industry average. It gives us a starting point for exploring the cost of heavy agent usage.

Using 20 working days per month, we can compare the subscription bill with published API prices.

ModelPer million uncached input tokensPer million output tokens
Claude Opus 5$5$25
GPT-5.6 Sol$4$20

Rates checked September 15, 2026: Anthropic pricing and Sol pricing. Sol's standard short-context rates are promotional through at least November 21, 2026.

ScenarioPer developer / month30 developers / month
Both subscriptions$400$12,000
Sol API rates$4,080$122,400
Opus API rates$5,100$153,000

Each API row shows what it would cost to use that model for all the estimated work. Using a mix of models would change the total.

The calculation assumes the output estimate includes reasoning tokens. The figures exclude caching costs, tool fees and taxes. Discounts or extra charges for very long requests could also affect the bill.

That is roughly 10–13 times the subscription bill at these model prices. It shows how differently the same usage can be priced, rather than what teams should expect to pay in future.

The gap can make subscriptions feel heavily subsidised. But API prices do not tell us what providers spend serving requests, so this comparison cannot show how much they earn or lose on a subscription.

Reporting on OpenAI's paid-product compute margin suggests that serving AI can carry substantial margins. That measure covers paid products broadly; it does not establish the API margin on these models or overall company profitability.

Competition and cheaper models

The comparison uses two premium models. Teams also have much cheaper options: Google launched Gemini 3.1 Flash-Lite at $0.25 per million input tokens and $1.50 per million output tokens. A lower-priced model will not suit every coding task, but teams can benefit wherever it delivers the quality they need.

That choice matters. As alternatives become more capable, Anthropic and OpenAI face pressure to keep their offerings attractive through price, performance and usage allowances. I expect that competition to benefit customers, even if providers respond in different ways.

There is also a history of falling costs. Stanford's 2025 AI Index reported a more than 280-fold price decline for GPT-3.5-level performance between November 2022 and October 2024. That does not predict future frontier-model prices, but better hardware and more efficient models could continue to make useful capability cheaper.

Pricing and allowances are worth keeping in mind as workflows grow. There is no need to assume a sudden price increase to make token economics relevant: cheaper models can encourage more usage too. If token prices halve while usage triples, the metered bill still rises by 50%.

Budget for useful work

As engineering teams become better at working with agents, token consumption is likely to keep growing. Cheaper models and stronger competition can make that growth more affordable, while opening up more work worth handing to AI.

For now, subscriptions can make extensive usage a relatively predictable expense. Over time, the choice of model, reasoning effort and number of parallel agents may become a more visible part of engineering budgets.

That leaves teams with a useful question: what are they getting back from the extra usage? Faster delivery, better software and more room to explore ideas can make it worthwhile. The token count alone cannot tell that story.

Happy coding!

Please share
𝕏finLINEtIw