The Transformation Brief

The AI Agent Cost Curve

Why the most expensive AI model isn't always the best business decision.

Ten generally available AI models shipped in the last seven weeks.

Claude Sonnet 5 arrived June 30, Grok 4.5 on July 8, GPT-5.6 in three variants on July 9, Kimi K3 on July 16, Gemini 3.6 Flash and 3.5 Flash-Lite on July 21, and Claude Opus 5 on July 24. Claude Fable 5 released June 9, went dark three days later under export controls, and has only been continuously purchasable since July 1. By the time most organizations finished evaluating one of them, another had claimed the top spot on a different benchmark.

I can prove that last sentence with a single leaderboard, because it happened to this newsletter while I was writing it. More on that below.

Every vendor will tell you this is a technology story, but what the release cadence actually creates is a governance problem: each launch promises more intelligence, lower latency, a bigger context window, or a lower price, and without a structured way to test those claims against your own work, procurement decisions get made on vendor positioning rather than measured outcomes.

Which is why every organization deploying AI at scale needs its own benchmarking strategy. Not to determine which model wins the internet's leaderboard, but to answer a far more practical question:

Which model delivers the best business outcome for our workflows at the lowest total cost?

Public benchmarks are a good starting point, but they cannot tell you how a model behaves inside your CRM, your ERP, your approval chain, or your security policy. The organizations that get the most out of AI over the next two years will be the ones with the most disciplined evaluation process, not the ones that adopt the newest model first.

The last time I priced a completed workflow

I have been on the other side of this exact evaluation problem, before AI was the thing being evaluated.

In 2019 I was the solutions engineer trying to get digital mortgage closing adopted at institutions like Chase, Bank of America, and Wells Fargo, and every evaluation started the same way: the bank graded the demo. Feature checklists, vendor comparisons, polished walkthroughs. Meanwhile the real question, what it costs to get one mortgage actually closed, start to finish, compliantly, sat unexamined in a spreadsheet nobody had built yet.

So I built it. Instead of defending features, I priced the completed workflow: what a paper closing cost the bank end to end, what a digital one would, and what the gap added up to at their volume. That model showed Chase they were leaving $45 million on the table annually, and it did what no demo had done in months of trying. It closed the deal, and adoption across our enterprise clients eventually went from near zero to 85 percent.

The lesson stuck with me because it keeps repeating: organizations evaluate technology on how impressive it looks in the room, when the only number that matters is the cost of finished work. Watching enterprises evaluate AI models on leaderboard position in 2026, I am seeing the same movie with better special effects.

Which brings me to the benchmark that caught my attention this month.

Intelligence is no longer the competitive advantage

Most AI benchmarks reward reasoning or conversational quality. Zapier's AutomationBench measures something executives actually care about: whether an agent can complete a real business workflow across Sales, Marketing, Operations, Support, Finance, and HR.

It was built from workflow patterns on a platform that processes more than two billion AI tasks a month across 3.7 million companies, which works out to roughly 540 automated tasks per company per month. This is not a lab exercise.

The benchmark drops an agent into a simulated business, hands it one instruction, and grades the world it leaves behind rather than the reasoning or the explanation along the way. Every assertion on the final state must pass or the task fails, and there is no partial credit in the headline number. Zapier's own framing is blunt: like production, mostly right is still wrong.

Businesses create value from completed work, not from eloquent responses. That shift, from measuring intelligence to measuring outcomes, is the most useful development in enterprise AI this year.

The question a CFO would ask

Most organizations still evaluate models with a consumer mindset, asking which one is smartest. A CFO would ask which one produces the lowest cost per completed workflow, so I ran that calculation on the current board.

I pulled the live leaderboard on July 30. Claude Opus 5, six days old, holds all five top positions, one for each of its effort settings, peaking at 26.2 percent of tasks completed correctly for $1.27 per task. Gemini 3.6 Flash on high effort sits sixth at 19.8 percent for 58 cents. GPT-5.6 Sol at maximum reasoning, which led this same board five days earlier, now ranks seventh at 18.1 percent for a dollar. Claude Fable 5 completes 17.4 percent at $2.03, the highest per-task price in the top ten.

Divide cost by completion rate, and the board inverts.

Gemini 3.6 Flash delivers a completed workflow for $2.93. The overall leader delivers one for $4.85, which puts the number-one model seventh out of ten on the only figure that appears on an invoice. And Fable 5, the most famous name on the board, delivers a completed workflow for $11.67, four times the price of a Gemini configuration that outscores it. Same benchmark, same six domains, and a four-times spread in what it costs to get one piece of work finished.

There is a second finding inside Opus 5's own five entries that is worth more than the ranking. Climbing the effort dial from low to max buys 3.8 points of completion for 69 percent more cost per task. And the high setting scores 23.0 percent for $1.03 while the medium setting scores 23.6 percent for 89 cents. With run-to-run variance typically inside one percent, that means 16 percent more spend bought a difference indistinguishable from noise, in the wrong direction, inside a single model's own dial.

The labs saw this coming. Anthropic shipped Opus 5 on July 24 with that explicit effort dial, priced at half of Fable 5 and marketed on efficiency rather than intelligence, after Fortune reported that enterprise customers had been complaining about Fable's token burn rate and the bills that came with it. The leaderboard now quantifies the complaint: Fable's per-task cost is the highest in the top ten, and the efficiency model that replaced it as the default delivers cheaper completed work at every one of its five settings. The interesting question has moved from what a model can do to what it costs to get finished work out of it.

The four costs most organizations ignore

Teams evaluating AI tend to look at token pricing, which is one line of four.

Inference cost. The published price to run the workflow. Across the current top ten, between 55 cents and $2.03 per task.

Retry cost. Cost divided by completion rate. This is where the 58-cent model beats the $1.27 leader by 40 percent per completed workflow, because you pay the token bill on every attempt and only some of them finish.

Failure cost. Zapier's analysis found that models declared success while actually failing in 72 percent of Claude Opus's failures, 91 percent of Gemini's, and 84 percent of GPT-5.4's. The agent reports the task done while the world state is wrong, and that cost never appears on an invoice. It appears as a bad record in a CRM, an approval in the wrong queue, or a customer told something untrue by a system nobody was auditing.

Human review cost. This one has a measurable proxy, and almost nobody uses it. Compare strict completion against partial completion. On strict scoring, Support and Finance sit near the bottom of the board together, but Artificial Analysis, scoring the same tasks at the objective level, found Support agents completing roughly 60 percent of individual objectives while Finance agents completed about a third. Similar headline scores, opposite operational verdicts: in Support the agent gets most of the way there and breaks one rule, so a reviewer can catch it cheaply, while in Finance the failure leaves so little to salvage that a reviewer is starting over.

The cheapest model becomes the most expensive when employees spend their afternoons correcting it, and the most expensive becomes economical when it finishes without intervention. Optimize for cost per completed workflow, not cost per token.

Different departments have different economics

The most useful thing about AutomationBench is that it reports by business domain rather than only in aggregate, which reflects how enterprises actually operate.

The domain spread on the current board is dramatic. Operations tops out at 56 percent while the best any configuration manages in HR is 13.3 percent, with Marketing at 29, Sales at 26.5, Support at 22, and Finance at 20.8 between them. The same portfolio of models is simultaneously half-reliable in one department and a one-in-seven proposition in another.

I wanted to know why, so I went into the public task files in Zapier's repository and read what the work actually is.

The Operations tasks are state synchronization between systems: update the incident ticket in Jira and document it in Confluence, archive the expired vendor contracts from Drive into Notion, move the maintenance queue from Sheets onto the ops board, put the emergency drill on the calendar and notify the channel. The environments are full of traps, stale rows and near-duplicate names and policies buried in old emails, but they are all the same species of trap: the truth is findable, and a diligent agent that looks harder wins. That is why models score 56 percent here. The work has a definite end state, the ground truth lives in the systems, and nobody's employment is involved.

The HR tasks are a different animal, and the task authors' own annotations tell the story. Raises over $15,000 require CFO processing, and HR cannot process them directly. Only the HRIS admin may touch the employee record. One departing employee has litigation pending, another is on garden leave, a third has a "legal hold, do not enroll" note buried in a spreadsheet column. My favorite is an offboarding task where the request cheerfully adds, "also go ahead and process their final severance payments while you're in there," which is precisely the thing the agent is not authorized to do. In HR, a large share of being correct is declining part of what you were asked, and the strict grader checks the refusals: who was not notified, which record was not touched. Agents are built to complete instructions, and this domain scores them on knowing when not to.

You cannot out-think an authority boundary the way you can out-search a stale row, which may be why HR is the one domain won by Opus 5 on low effort, the cheapest setting it has, at 75 cents a task. More reasoning finds more data, but it does not reliably produce more restraint. Banks understood this about wire approvals decades ago, and it is the same reason my 2019 mortgage closings kept humans on the compliance gates.

That reading gives you a practical screen for your own backlog. A workflow is automatable today if you could write down its finished state as a checklist, the facts the agent needs live somewhere in your systems, and being correct never depends on who is allowed to act. Workflows that fail the third test are not model-selection problems, and no amount of effort dial will make them one.

Then look at which configuration wins each domain, because it is frequently not the expensive one.

In Operations, Opus 5 at maximum effort leads at 56.0 percent with Gemini 3.6 Flash on high at 55.0, a gap inside the variance band, and Gemini's per-task cost is less than half of the leader's. In Finance they tie outright at 20.8 percent, $1.27 against 58 cents. Run a thousand finance tasks a month and the tie choice is the difference between $1,270 and $580 for the same 208 completions, which is $8,280 a year for identical measured performance, earned by reading one column further.

Stranger still: in three of the six domains, the winning configuration of the winning model is not its maximum effort. Support is won by Opus 5 on medium, Marketing on xhigh, and HR on low. More thinking does not reliably buy more finished work, and in some departments it measurably buys less.

Build an AI portfolio, not an AI monopoly

The biggest misconception in enterprise AI is that everyone in the company should use the same model. Organizations already run multiple cloud providers, multiple security platforms, and multiple data tools, and there is now hard evidence for treating AI the same way.

Zapier's researchers compared which specific tasks each model passed, and their two top performers overlapped by a Jaccard similarity of 0.17. Seventy-one percent of the tasks one model solved, the other failed, and seventy percent of the second model's passes were missed by the first. Two models with nearly identical overall scores turned out to hold almost entirely different competence.

Read that carefully, because it kills a common argument. If model capabilities were converging, standardizing on one vendor would be the rational move and this newsletter would be arguing against itself. They are not converging but diverging in which work they can actually finish. A single standard model across six departments is the wrong shape for that distribution, and a portfolio matched to workflow economics is the right one.

Two moves that beat model selection entirely

Before you shop, two levers sit upstream of the buying decision, and both cost nothing.

The first is your tool surface. Zapier ran the same models against three different tool configurations. Gemini 3.1 Pro passed 9.6 percent when it had to discover raw API endpoints, 12.8 percent with a curated toolset, and 14.3 percent with a toolset scoped to the task, while Haiku 4.5 went from 1.5 percent to 3.8 percent across the same range. Same model, same token spend, and the largest single improvement in the entire dataset. Most organizations buy a better model when what they needed was a narrower tool list.

The second is your own noise floor. Zapier holds variance inside one percent across a 600-plus task evaluation set, while on a 20-task internal evaluation, one task flipping moves your score five points, which is five times their margin. Anything under roughly 30 tasks per domain is a smoke test, not a procurement input, and without that number you cannot tell whether an effort tier is buying performance or selling you variance. The Opus 5 medium-versus-high finding above is exactly the kind of thing you cannot see without it.

The board moved while I was writing this

Now the part I promised in the opening.

I first pulled this leaderboard on July 25. GPT-5.6 Sol at maximum reasoning led it at 18.1 percent, Fable 5 sat second, and Opus 5, Gemini 3.6 Flash, and Kimi K3 appeared nowhere on the board.

Five days later, Sol posts the identical 18.1 percent and ranks seventh. Opus 5 arrived and took the top five slots outright. Gemini 3.6 Flash landed in sixth and eighth and quietly became the best value on the board. Every ranking-based conclusion from my first pull was obsolete inside a week, while the scores themselves never changed. Nothing got worse. The market simply got measured faster than any procurement memo could circulate.

That gap deserves a name. Measurement debt is the distance between the models available to your organization and the models you have actually tested against your own work, and it behaves like technical debt in every respect but one: almost nobody tracks it. Zapier, whose full-time job this is, still has no score for Kimi K3 two weeks after launch. Your internal evaluation backlog is longer than theirs.

I would go further. Measurement debt is the newest room in a house I have spent 20 years working in, the one I call the AI Adoption Gap: the space between a signed AI contract and actual enterprise transformation, where the sale, the implementation, and the change management belong to three different teams and the outcomes belong to nobody. Buying capability you cannot measure is how organizations fall into that gap with a receipt in hand. Every purchasing decision made inside it is a guess wearing a spreadsheet.

The debt is why an AI portfolio needs continuous evaluation rather than a one-time vendor bake-off. A benchmark you ran in April is describing a market that no longer exists. Mine barely survived five days. The organizations that win on AI will be the ones that can measure cost per completed workflow inside their own systems, and are still measuring it next quarter.

Three questions I keep turning over after writing this:

  1. If someone asked for your cost per completed workflow tomorrow morning, who in your organization could produce the number, and how old would it be?
  2. Which of your AI wins has actually been audited for what changed in your systems, rather than for what the agent reported about itself?
  3. When your team picked its current model, was that a measurement, or a moment, and would anyone be able to tell you which?

Leaderboard figures pulled July 30, 2026, directly from Zapier's live AutomationBench leaderboard (v1.0.5). The July 25 comparison reflects the board's composition before Opus 5 and Gemini 3.6 Flash were added. Silent-failure, task-overlap, and tool-configuration figures are from the AutomationBench paper (arXiv:2604.18934). Objective-level domain scoring is from Artificial Analysis's AutomationBench-AA. Opus 5 pricing and enterprise reaction are from Fortune's July 24 coverage. Cost per completed workflow is my calculation from Zapier's published score and cost figures, not a metric Zapier reports. The domain-difficulty reading is my analysis of the public task definitions in Zapier's GitHub repository, not Zapier's claims.