Choosing the metric

Which number gets reported, and whether it tracks anything the business feels.

7documents on this topic
7organizations represented
2issues named
6sourced citations
3sourced statistics

The state of it

One of 6 topics within Evaluation & observability.

7 documents from 7 organizations address which number gets reported. It is the thinnest topic in the research, which is itself worth noting given that choosing the wrong metric invalidates everything measured downstream of it.

The pattern visible across the wider set is that the reported number is the one that is easy to collect - usage, seats, calls, use-case counts - rather than the one the business feels. BCG's split is the sharpest illustration available: 82% of CEOs more optimistic about AI ROI than a year earlier, against only 5% of companies generating value at scale and 60% reporting no material gains. Something is being measured enthusiastically, and it is not the thing that shows up in results.

The issues, by agreement

How many independent organizations name each issue as a problem. An issue is only as real as the number of separate publishers that identify it, so the count is the ranking. Bars are organizations, not documents. Where the count reads ours, no publisher here states the issue and the analysis is our own.

Where they disagree

No contradictions recorded on this topic yet.

The issues in full

Each issue carries the organizations that name it, the numbers behind it, and the remedies proposed - with the concrete steps under each. Every citation points at a section of a named document, so any count here can be checked.

Issue 012 organizations name it2026 evidence

The reported number is the one that is easy to collect

Usage, seats, calls and use-case counts are instrumented by default and reported upward because they exist. None of them is felt by the business.

82%CEOs more optimistic about AI ROI than a year earlierBoston Consulting Group · Jan 2026
5%Companies generating AI value at scaleBoston Consulting Group · Jul 2026
60%Companies reporting no material gains despite substantial investmentBoston Consulting Group · Jul 2026
How to fix it — 1 approach, 3 steps

Report one metric the business already feels

Pick a number the operating review discussed before AI existed - cycle time, cost to serve, resolution rate - and report AI against that.

Done when Each deployment reports against a business metric the operating review discussed before AI existed, usage and seat counts no longer go upward, and a metric that did not move is reported as not having moved.

  1. Choose one pre-existing business metric per deployment and report against it.0-30 daysCOO
  2. Stop reporting usage, seats and use-case counts upward.30-90 daysCIO
  3. Where the felt metric did not move, say so rather than substituting an activity measure.ongoingCFO
The evidence — 3 documents
OrganizationDocumentPosition
Boston Consulting GroupConsultancy · July 2026Driving cost advantage with AIOur reading Reports 82% of CEOs more optimistic about AI ROI than a year earlier against only 5% of companies generating AI value at scale and 60% reporting no material gains despite substantial investment.Executive summarynames it
USTConsultancy · November 2025CIOs guide to agentic AIOur reading Specifies outcome-based metrics as a requirement, and treats measurable business impact as the test of an agentic deployment.Outcome-based metricsnames it
OpenAIFrontier lab · August 2026A scorecard for the AI ageOur reading Proposes measuring the full cost of producing a successful outcome against the value that outcome creates, rather than any activity or unit-price measure.Cost per successful taskproposes a fix

Issue 021 organization name it1 qualifies itnewest evidence Jun 2026

One reliability bar is set for the whole estate, when the bar is a function of what fails

Reliability is discussed as a property of the technology - is it good enough yet - rather than as a property of each use. The same error rate is either unremarkable or disqualifying depending entirely on what the system is attached to.

Palantir puts the arithmetic plainly: an agent that fails one time in a hundred is unremarkable drafting sales emails and unacceptable shipping code to production, and the same gap holds in defence, financial services and much of healthcare. The consequence is that a single organisation-wide reliability standard is either too strict to permit anything useful or too loose to be safe, and usually both at once in different parts of the estate. The harder half is that the number is not the real problem: not knowing WHEN the system will fail is worse than knowing it fails one per cent of the time, because an unpredictable failure cannot be designed around.

How to fix it — 2 approaches, 5 steps

Set the bar per use, from the consequence of being wrong

Classify each deployment by what happens when it fails, and attach a different reliability standard to each class. One number for everything is a number nobody can act on.

Done when Every deployed use records what a single wrong output would cost and sits in a tier with its own error ceiling and its own approval to run, and reporting shows predictability of failure as well as its rate.

  1. List every deployed use and record what a single wrong output would cost.0-30 daysRisk
  2. Assign each to a tier with its own error ceiling and its own approval to run.30-90 daysRisk
  3. Measure predictability of failure, not only its rate, and report both.90-180 daysCTO

Treat unpredictable failure as disqualifying

Where the system cannot say when it is likely to be wrong, keep it away from anything irreversible, whatever its average accuracy.

Done when No unattended deployment runs without a confidence or abstention signal, and the default route for a low-confidence case is to a person rather than the exception.

  1. Require a confidence or abstention signal before any unattended deployment.30-90 daysCTO
  2. Route low-confidence cases to a person by default rather than by exception.30-90 daysCOO
The evidence — 3 documents
OrganizationDocumentPosition
PalantirEnterprise · June 2026Palantir: Governing AI agentsOur reading States the case directly: the tolerable error rate is set by what the agent is wired to, not by the agent, and recent capability gains have bought only small improvements in reliability.Reliability in enterprise contextnames it
AnthropicFrontier lab · December 2024Building effective agentsOur reading Qualifies it from the design side: the case for an agent rests on whether the path and the number of steps can be known in advance, and the autonomy it buys is conditioned per deployment on a trusted environment and errors that stay containable rather than compound.Where agents fitqualifies it
USTConsultancy · November 2025CIOs guide to agentic AIOur reading Proposes different numeric bars for different settings - a high completion rate for routine service work, a far stricter failure ceiling in regulated ones - which is this issue answered rather than argued.Task completion and severity thresholdsproposes a fix

Who is represented

This dossier is drawn from 10 organizations working on the subject, 5 of which are cited directly in the issues above.

Consultancy — 4

Boston Consulting Group 1 UST 1 Capgemini 1 Deloitte 1

Institution — 1

World Economic Forum 1

Hyperscaler — 2

Google Cloud 1 Microsoft 1

Frontier lab — 2

Anthropic 2 OpenAI 1

Enterprise — 1

Palantir 1