Total cost of ownership

The denominator: run cost, tokens, rework, and the human review nobody budgeted.

30documents on this topic
21organizations represented
2issues named
5sourced citations
1sourced statistics

The state of it

One of 7 topics within ROI measurement.

30 documents from 21 organizations address what AI actually costs to run. The material has moved noticeably in the last year, from token prices toward total cost of ownership - and the reason is that unit prices kept falling while bills kept rising.

The consistent theme is that the denominator in most ROI calculations is wrong, because it contains the model and not the work around it: retries when a cheaper model fails, human review of output, rework, the orchestration layer, and the governance tooling. EY's framing is the useful one - the lowest token price does not produce the lowest total cost, because a cheaper model that fails and retries costs more than an expensive one that succeeds first time.

The issues, by agreement

How many independent organizations name each issue as a problem. An issue is only as real as the number of separate publishers that identify it, so the count is the ranking. Bars are organizations, not documents. Where the count reads ours, no publisher here states the issue and the analysis is our own.

Where they disagree

No contradictions recorded on this topic yet.

The issues in full

Each issue carries the organizations that name it, the numbers behind it, and the remedies proposed - with the concrete steps under each. Every citation points at a section of a named document, so any count here can be checked.

Issue 012 organizations name it2026 evidence

The business case funds the build and not the running

Monitoring, retraining, evaluation and the orchestration layer are continuing costs that appear in no approval paper, so they compete with new work and lose.

How to fix it — 1 approach, 3 steps

Require a run-cost line before approval

Every approval carries a three-year run cost with a named owner, or it does not get approved.

Done when Every approval carries a three-year run cost covering inference, monitoring, evaluation and retraining, assigned to a named operating budget owner at approval, with actual reviewed against estimate at twelve months.

  1. Add a mandatory three-year run-cost estimate covering inference, monitoring, evaluation and retraining.0-30 daysCFO
  2. Assign the run cost to an operating budget owner at the point of approval, not at go-live.0-30 daysCFO
  3. Review actual against estimated run cost at twelve months.90-180 daysCFO
The evidence — 3 documents
OrganizationDocumentPosition
EYConsultancy · June 2026Unlocking agentic value: a new investment disciplineOur reading Places escalating cost first among the reasons agentic projects are forecast to be cancelled.Escalating costs as a cancellation drivernames it
McKinsey & CompanyConsultancy · March 2026Recalibrating technology budgets for the AI eraOur reading Identifies the balance between run and change spend as the core CIO budget challenge, which is exactly where unbudgeted AI operating costs land.Run spend versus change spendnames it
KPMGConsultancy · April 2026Building the AI business caseOur reading Sets out an investment assessment covering hardware, software and people rather than licence cost alone.Building the AI business caseproposes a fix

Issue 021 organization name it2026 evidence

The cost calculation leaves out the human review that makes the output usable

Cost per token, and even cost per call, omits the reviewer. In workflows where output is checked before use, that reviewer is often the largest single cost.

40%Agentic AI projects forecast to be cancelled by end of 2027EY · Jan 2025
How to fix it — 1 approach, 3 steps

Price the accepted outcome, not the call

Divide total cost - model, tools, retries, and human review time - by the number of outputs that met the bar. That is the only figure that compares across models.

Done when An accepted outcome is defined for the workflow, retries and human review minutes are captured alongside token spend, and models are compared on cost per accepted outcome rather than unit price.

  1. Define what counts as an accepted outcome for the workflow before measuring anything.0-30 daysCOO
  2. Capture retries and review minutes alongside token spend.30-90 daysCIO
  3. Compare models on cost per accepted outcome, never on unit price.ongoingCIO
The evidence — 2 documents
OrganizationDocumentPosition
EYConsultancy · June 2026Unlocking agentic value: a new investment disciplineOur reading Argues for an investment discipline built on total escalating cost rather than unit price, relaying Gartner's forecast that more than 40% of agentic projects are abandoned before 2028 over cost, unproven value or weak controls.A new investment disciplinenames it
McKinsey & CompanyConsultancy · March 2026Recalibrating technology budgets for the AI eraOur reading Frames the core budget challenge as the balance between run spend and change spend, which is where an unbudgeted run cost surfaces.Balancing run spendproposes a fix

Who is represented

This dossier is drawn from 22 organizations working on the subject, 3 of which are cited directly in the issues above.

Consultancy — 8

McKinsey & Company 4 EY 1 KPMG 1 Boston Consulting Group 2 Cognizant 1 Deloitte 1 Genpact 1 UST 1

Institution — 5

FinOps Foundation 2 Australian Institute of Company Directors 1 Cloud Security Alliance 1 Thomson Reuters 1 World Economic Forum 1

Hyperscaler — 2

Google Cloud 3 Microsoft 2

Frontier lab — 2

Anthropic 1 OpenAI 1

Vendor — 4

IntuitionLabs 2 Cathay Capital 1 Predibase 1 Writer 1

Other — 1

CISA 1