Baseline & attribution

Whether the gain can be shown to have come from the AI rather than from everything else that changed.

23documents on this topic
17organizations represented
2issues named
3sourced citations
0sourced statistics

The state of it

One of 7 topics within ROI measurement.

23 documents from 17 organizations touch on how a gain is established. The gap between the academic and the commercial material is at its widest here.

The controlled studies do the thing that makes a number mean something: they compare against a group that did not get the tool. Almost nothing else in this set does. What the commercial material generally offers is a before-and-after on a process that was also being redesigned, staffed differently, and measured with more attention than it had ever received - all at the same time.

That does not make the numbers wrong. It makes them unattributable, which is a different problem and a more awkward one to raise after the investment is approved.

The issues, by agreement

How many independent organizations name each issue as a problem. An issue is only as real as the number of separate publishers that identify it, so the count is the ranking. Bars are organizations, not documents. Where the count reads ours, no publisher here states the issue and the analysis is our own.

Where they disagree

No contradictions recorded on this topic yet.

The issues in full

Each issue carries the organizations that name it, the numbers behind it, and the remedies proposed - with the concrete steps under each. Every citation points at a section of a named document, so any count here can be checked.

Issue 01Our analysisnewest evidence Apr 2023

Almost nothing outside the academic work has a control group

A before-and-after comparison on a process that changed in several ways at once cannot isolate the contribution of the AI. Most reported enterprise ROI is of this kind.

The three controlled studies here are useful precisely because they are boring in design: randomise, withhold the tool from some, compare. That design is what lets Brynjolfsson and colleagues say the 15% is caused by the assistant rather than by the attention that came with the rollout. Enterprises rarely do this, usually for good operational reasons - but the consequence is that the resulting figure cannot survive a serious challenge, and finance functions increasingly know it.

How to fix it — 2 approaches, 5 steps

Withhold the tool from one comparable group

A staged rollout is already a natural experiment. Sequence it so the last group is a usable comparison rather than just the last group.

Done when Two comparable teams were measured on the same metric over the same window with one held back for a defined period, and the report states the difference between them rather than the improvement in the deployed one.

  1. Choose two comparable teams and deploy to one first, holding the second for a defined period.0-30 daysCOO
  2. Measure both on the same metric over the same window.30-90 daysCOO
  3. Report the difference, not the improvement.30-90 daysCFO

Freeze the metric definition before deployment

A measure that is refined during the rollout will show improvement whether or not anything improved.

Done when The metric definition, its data source and its owner were written down before deployment, and any change since carries a recorded reason and a restated history.

  1. Write the metric definition, the data source and the owner before anything is deployed.0-30 daysCFO
  2. Change the definition only with a recorded reason, and restate history when you do.ongoingCFO
The evidence — 2 documents
OrganizationDocumentPosition
arXiv (research)Academic · February 2023The Impact of AI on Developer Productivity: Evidence from GitHub CopilotOur reading Randomly assigns recruited developers to treatment and control on an identical task, isolating the tool as the only difference.Experimental designproposes a fix
National Bureau of Economic ResearchAcademic · April 2023Generative AI at WorkOur reading Uses a staggered introduction across sites so that untreated agents serve as the comparison, which is what allows the productivity change to be attributed to the assistant.RCT analysis; staggered rollout designproposes a fix

Issue 02Our analysisnewest evidence Jun 2025

The process was never measured before the AI arrived

Many AI investments target processes that had no baseline at all, so the improvement is being compared to an estimate made after the fact.

How to fix it — 1 approach, 3 steps

No baseline, no funding

Make a measured pre-AI baseline a mandatory field in the investment paper, and reject papers without one at least once, visibly.

Done when The investment template requires the baseline, who measured it and when, at least one paper has been visibly rejected for lacking it, and value is reported against the original rather than a restated figure.

  1. Add baseline measured, by whom, on what date, as a required field in the investment template.0-30 daysCFO
  2. Fund a two-week measurement exercise where no baseline exists - it is cheaper than the argument later.30-90 daysCFO
  3. Report all value against the original baseline, never a restated one.ongoingCFO
The evidence — 1 document
OrganizationDocumentPosition
CapgeminiConsultancy · June 2025The blueprint to scaling AI for business transformationOur reading Proposes assessing how amenable each area is before committing, which requires a measured current state.Assessing amenability and building a roadmapproposes a fix

Who is represented

This dossier is drawn from 18 organizations working on the subject, 3 of which are cited directly in the issues above.

Consultancy — 6

Capgemini 1 Boston Consulting Group 2 PwC 2 Accenture 1 EY 1 Infosys 1

Institution — 3

Cloud Security Alliance 2 NIST 2 World Economic Forum 1

Academic — 3

arXiv (research) 1 National Bureau of Economic Research 1 Carnegie Mellon SEI 1

Hyperscaler — 3

Google Cloud 3 IBM 1 Microsoft 1

Frontier lab — 1

Anthropic 1

Vendor — 1

Palo Alto Networks 1

Other — 1

CISA 1