What the evidence actually shows

Measured effects on real work, who they accrue to, and where they reverse.

45documents on this topic
20organizations represented
3issues named
7sourced citations
6sourced statistics

The state of it

One of 7 topics within ROI measurement.

45 documents from 20 organizations report what AI does to output. Almost all of them report a percentage. Very few of them ran a control group.

That distinction turns out to matter more than any single number - and it comes with a caveat worth stating first. The only controlled studies available are from 2023, and they measured systems a generation behind current models. Their percentages are not current and are not presented here as statistics. Their shape is what transfers, and no equivalent controlled study has been published on the models being deployed now. Every confident percentage in the newer material rests on no control group.

What the controlled work found is that the average gain is real and, more importantly, that it is not evenly distributed. Brynjolfsson, Li and Raymond, working with 5,172 customer support agents, found a 15% average lift in issues resolved per hour - and found that it went almost entirely to the least experienced workers, with the most skilled seeing small speed gains and a small decline in quality. Peng and colleagues found the same shape in a developer experiment: a large average effect, concentrated among the less experienced.

The consulting material reports function-level percentages that read as uniform across a workforce. On the evidence here they are not, and the difference changes who you deploy to first.

The issues, by agreement

How many independent organizations name each issue as a problem. An issue is only as real as the number of separate publishers that identify it, so the count is the ranking. Bars are organizations, not documents. Where the count reads ours, no publisher here states the issue and the analysis is our own.

Who takes which position

The chart above counts positions; this shows whose they are. Read down a column for what one organization holds across the whole topic, and across a row for who lines up on one issue. Where a cell carries more than one position, the strongest is shown and the rest are in the tooltip.

Ddisputes it Qqualifies it Nnames it as a problem Pproposes a fix
What the evidence actually shows: 3 issues against the 5 organizations cited on them. The number under each name is how many of these issues it is cited on.
Issue Boston Consulting Group · 1 IBM · 1 KPMG · 1 National Bureau of Economic Research · 1 arXiv (research) · 1
A faster task is not a faster process N · · · Q
Function-level percentages describe potential, and are read as results · Q N · ·
The measured gains go to the least experienced, and can reverse for your best people · · · N ·

A dot means this organization is not cited on that issue. It does not mean they are silent on it: an organization is cited where its document takes a position we could locate, and the absence of a citation is the absence of a finding, not a finding of absence. Who is represented lists everyone working on this topic, including those not cited above.

Where they disagree

No contradictions recorded on this topic yet.

The issues in full

Each issue carries the organizations that name it, the numbers behind it, and the remedies proposed - with the concrete steps under each. Every citation points at a section of a named document, so any count here can be checked.

Issue 011 organization name it1 qualifies itnewest evidence Jul 2026

A faster task is not a faster process

Most reported gains are measured at the level of an isolated task. Whether the process around it got faster is a different measurement, and usually an unmade one.

How to fix it — 1 approach, 3 steps

Measure the cycle, not the step

Pick the end-to-end cycle time the business already cares about and measure that. If it did not move, the task-level gain went somewhere else.

Done when The end-to-end metric was named before deployment and measured for one full cycle before and after on the same definition, and where the step moved but the cycle did not, the queue between steps is recorded as the constraint.

  1. Name the end-to-end metric before deployment: lead time, cycle time, cost to serve.0-30 daysCOO
  2. Measure it for one full cycle before and after, on the same definition.30-90 daysCOO
  3. Where the step improved and the cycle did not, treat the queue between steps as the real constraint.90-180 daysCOO
The evidence — 3 documents
OrganizationDocumentPosition
Boston Consulting GroupConsultancy · July 2026Driving cost advantage with AIOur reading Lists this second among five traps it holds stop most companies turning AI into durable cost advantage: speed applied inside a process nobody redesigned first lifts volume while leaving the cost base where it was, which is offered as one reason so many report no material gain.The five traps: linking AI adoption to cost reductionnames it
arXiv (research)Academic · February 2023The Impact of AI on Developer Productivity: Evidence from GitHub CopilotOur reading Measures a single defined programming task completed as quickly as possible, which isolates the effect cleanly and by construction says nothing about the surrounding workflow - a limit the paper states of itself.Experimental designqualifies it
Boston Consulting GroupConsultancy · December 2024AI value creation in leading enterprisesOur reading Places the large majority of the required effort in people and process change rather than in the technology, which is where task-level gains either become throughput or do not.The 10-20-70 framingproposes a fix

Issue 021 organization name it1 qualifies itnewest evidence Mar 2025

Function-level percentages describe potential, and are read as results

The widely cited function productivity figures are assessments of what is technically feasible, usually assuming redesigned processes. They circulate as if they were measured outcomes.

18%Addressable GenAI opportunity, upper boundKPMG · Jan 2025
82%CEOs more optimistic about AI ROI than a year earlierBoston Consulting Group · Jan 2026
5%Companies generating AI value at scaleBoston Consulting Group · Jul 2026
60%Companies reporting no material gains despite substantial investmentBoston Consulting Group · Jul 2026
How to fix it — 1 approach, 3 steps

Label every number as potential or observed

One word in front of each figure - modelled, or measured - prevents most of the confusion in an investment paper.

Done when Every productivity figure in board and investment material is tagged modelled or measured, each modelled one names the assumption it rests on, and the two are tracked separately so the gap is visible.

  1. Tag every productivity figure in board and investment material as modelled potential or observed result.0-30 daysCFO
  2. Require any modelled figure to name the assumption it depends on, usually process redesign.0-30 daysCFO
  3. Track the two separately over time so the gap between them is visible.ongoingCFO
The evidence — 2 documents
OrganizationDocumentPosition
KPMGConsultancy · March 2025Value at StakeOur reading Sizes the addressable opportunity at 4-18% of EBITDA annually depending on sector, and states directly that executives struggle to quantify expected returns and set realistic targets.The GenAI opportunitynames it
IBMHyperscaler · November 2024ROI of AIOur reading Puts the figure on self-reported ground: technology buyers rating their own returns on an unweighted online panel, with a share saying outright that they find the return difficult to measure at all.Survey methodologyqualifies it

Issue 031 organization name itnewest evidence Apr 2023

The measured gains go to the least experienced, and can reverse for your best people

Average productivity numbers hide a distribution. In the controlled studies the lift is large for novices, small for experts, and slightly negative on quality for the most skilled - which is the opposite of how most deployment plans are sequenced.

Read the shape of this finding, not its size. Both controlled studies here are from 2023 and measured systems a generation behind anything you would deploy now, so their headline percentages are not current and are deliberately not shown as statistics on this page. What transfers is the distribution: Brynjolfsson, Li and Raymond found treated agents with two months of tenure performing as well as untreated agents with more than six, and found the most skilled seeing small speed gains alongside a small decline in quality. Peng and colleagues found the same direction in software development. The uncomfortable part is what has happened since: capability has moved several generations and no equivalent controlled study has been published on current models. Every confident percentage in the more recent material rests on no control group at all.

30%Productivity gain reported as feasible today across functionsBoston Consulting Group · Sep 2025
50%Productivity gain reported as reachable with reshaped processesBoston Consulting Group · Sep 2025
How to fix it — 2 approaches, 5 steps

Measure the gain by tenure band, not as an average

Report the effect separately for your least and most experienced cohorts. An average across both will mislead every deployment decision downstream.

Done when Every productivity figure reaching the investment committee is split into at least three tenure or skill bands and reported as a distribution rather than a mean, with a quality reading for the top band beside it.

  1. Split the measured population into at least three tenure or skill bands before reporting any productivity figure.0-30 daysCHRO
  2. Report the distribution, not the mean, to the investment committee.30-90 daysCFO
  3. Watch the top band specifically for quality regression, which the average will hide.ongoingCOO

Sequence deployment toward the least experienced

If the evidence says the lift concentrates among newer workers, pilot there - not with the strongest team, where it will look least impressive.

Done when The first production deployment sits in the function with the largest low-tenure population and the highest onboarding cost, and time-to-competence is measured there rather than throughput alone.

  1. Identify the functions with the largest population of low-tenure staff and the highest onboarding cost.0-30 daysCHRO
  2. Run the first production deployment there and measure time-to-competence, not just throughput.30-90 daysCOO
The evidence — 2 documents
OrganizationDocumentPosition
National Bureau of Economic ResearchAcademic · April 2023Generative AI at WorkOur reading Data from 5,172 customer support agents: 15% average increase in issues resolved per hour, with less experienced and lower-skilled workers improving both speed and quality while the most experienced and highest-skilled see small gains in speed and small declines in quality.Abstract; Section 4.1 heterogeneity (2023 study, GPT-3.5-era assistant)names it
National Bureau of Economic ResearchAcademic · April 2023Generative AI at WorkOur reading Treated agents with two months of tenure perform as well as untreated agents with more than six months, indicating the tool compresses time-to-competence rather than lifting everyone equally.Experience curve findingsnames it

Who is represented

This dossier is drawn from 20 organizations working on the subject, 5 of which are cited directly in the issues above.

Consultancy — 11

Boston Consulting Group 14 KPMG 2 Capgemini 4 McKinsey & Company 3 Accenture 2 Arthur D. Little 1 Deloitte 1 EY 1 Genpact 1 PwC 1 UST 1

Institution — 1

World Economic Forum 1

Academic — 2

arXiv (research) 2 National Bureau of Economic Research 1

Hyperscaler — 4

IBM 4 Microsoft 2 Databricks 1 Google Cloud 1

Frontier lab — 1

Anthropic 2

Vendor — 1

TechWolf 1