Where a number came from, and whether two systems agree on it.
35documents on this topic
23organizations represented
3issues named
5sourced citations
0sourced statistics
The state of it
One of 5 topics within Data readiness.
Lineage is usually implemented at ingestion and lost everywhere after it. That is survivable in a reporting stack, where a human reads the number and knows roughly where it came from. It is not survivable when a model composes an answer from four retrieved fragments and presents it in a sentence.
The requirement that falls out of the engineering accounts is specific: lineage has to travel through extraction, chunking, enrichment and indexing, not merely be recorded at the door. A passage retrieved from a vector index should be traceable to the document, the version and the section it came from - otherwise a wrong answer cannot be diagnosed, a withdrawn document cannot be purged from what the system will still say, and an assurance question about where a figure came from has no answer.
This is also where the golden source problem resurfaces in a new form. Two systems disagreeing about a customer count was a reconciliation meeting. Two retrieval paths disagreeing is a system that gives different answers to the same question depending on which fragment ranked higher that day.
The issues, by agreement
How many independent organizations name each issue as a problem. An issue is only as real as the number of separate publishers that identify it, so the count is the ranking. Bars are organizations, not documents. Where the count reads ours, no publisher here states the issue and the analysis is our own.
The chart above counts positions; this shows whose they are. Read down a column for what one organization holds across the whole topic, and across a row for who lines up on one issue. Where a cell carries more than one position, the strongest is shown and the rest are in the tooltip.
Ddisputes itQqualifies itNnames it as a problemPproposes a fix
Lineage, provenance & the golden source: 3 issues against the 3 organizations cited on them. The number under each name is how many of these issues it is cited on.
A dot means this organization is not cited on that issue. It does not mean they are silent on it: an organization is cited where its document takes a position we could locate, and the absence of a citation is the absence of a finding, not a finding of absence. Who is represented lists everyone working on this topic, including those not cited above.
Where they disagree
No contradictions recorded on this topic yet.
The issues in full
Each issue carries the organizations that name it, the numbers behind it, and the remedies proposed - with the concrete steps under each. Every citation points at a section of a named document, so any count here can be checked.
Issue 012 organizations name it2026 evidence
In a chain of agents, a weak signal upstream does not average out
Data quality is assessed as a property of a dataset, on the assumption that errors are noise. In a chain where each step consumes the previous step, a single weak input upstream compounds into the output rather than diluting.
This is the specific reason data debt tightened when agents arrived, and it is the argument for spending on quality at the head of a chain rather than evenly across the estate. Each step is only ever as sound as whatever the step before handed it; the failure is not that one step is wrong but that nothing after it can tell, because each subsequent step treats its input as given. The published figure attached to this is broad but useful: data quality failures leave 42% of analytics and AI initiatives delayed, underperforming or failed - a range, not a failure rate, and worth quoting as the range.
Identify the first dataset in each agent chain and hold it to a higher standard than the ones downstream, which inherit whatever it carries.
Done when Each live agent chain names the dataset its first step consumes, those datasets are held to a written standard higher than the rest of the estate, and one deliberate degradation test is on record with what downstream noticed.
For each live agent chain, name the dataset the first step consumes.0-30 daysHead of architecture
Set a higher quality bar for those head datasets than for the rest of the estate.30-90 daysChief data officer
Test one chain by degrading the head input deliberately, and see whether anything downstream notices.90-180 daysHead of data engineering
The evidence — 3 documents
Organization
Document
Position
GenpactConsultancy · June 2026
$18 trillion in trapped AI valueOur reading Reports that agentic chains pass data quality forward step by step, so one weak signal near the start cascades into errors that compound further down, and that quality failures leave 42% of analytics and AI initiatives delayed, underperforming or failed outright.Compounding rather than averaging
names it
UberEnterprise
Enhanced Agentic-RAG: chatbots at near-human precisionOur reading Names it without caveat: when the wrong passage comes back the model produces errors regardless, and the reported gain - a quarter more acceptable answers and a large fall in wrong advice - came from repairing document processing and retrieval, upstream of the model that writes the answer.Where precision is won or lost
names it
McKinsey & CompanyConsultancy · June 2026
Data readiness for scaling AI impactOur reading Proposes carrying lineage through each transformation, which is what allows a compounded error to be walked back to the step that introduced it.Traceability through the pipeline
proposes a fix
Issue 02Our analysis2026 evidence
Lineage is recorded at the door and lost through the pipeline
The catalogue knows where a file came from. Nothing knows which version of which section produced the passage the model just used to answer a customer.
The practical tests are simple and most estates fail them: can you take a sentence the system produced and name the document, version and section behind it; and if a document is withdrawn today, can you be sure the system will stop repeating it. Both require lineage to survive extraction, chunking and indexing rather than being an ingestion-time record. Retrofitting it means rebuilding the index, which is why it belongs in the pipeline pattern from the start.
Require any produced sentence to be traceable to a source section
Make traceability a test rather than an aspiration: take an answer the system gave and walk it back to the document, version and section.
Done when Three answers a live system produced have each been traced to a document, version and section, and a withdrawal test shows the system stopped repeating a document after it was removed.
Take three answers a live system produced and try to trace each to a source section.0-30 daysHead of data engineering
Carry document, version and section through extraction, chunking and indexing.30-90 daysHead of data engineering
Test withdrawal: remove a document and confirm the system stops repeating it.90-180 daysHead of risk
The evidence — 1 document
Organization
Document
Position
McKinsey & CompanyConsultancy · June 2026
Data readiness for scaling AI impactOur reading Describes preserving meaning, lineage and control as content moves through ingestion, extraction, quality checks, metadata and indexing, rather than recording origin once at the start.Lineage preserved through the pipeline
proposes a fix
Issue 03Our analysis2026 evidence
Two retrieval paths give two answers and neither is wrong
The same question asked twice returns different answers, because different fragments ranked higher. There is no golden source for an answer composed at query time.
The old version of this problem was two dashboards disagreeing, and the fix was a master record. The new version is harder because the answer does not exist until it is asked for, and it is assembled from whatever the index returned. That makes the discipline more important rather than less: a designated authoritative source per domain, superseded documents removed from the index rather than left to rank low, and a rule about which source wins when two disagree - decided in advance, not by ranking.
Name the authoritative source per domain, and retire the rest from the index
Decide in advance which source wins for each domain, and remove superseded material from the index rather than trusting ranking to bury it.
Done when Each domain the system answers on names its authoritative source in writing, and superseded documents have been removed from the index rather than left to rank low.
For each domain a system answers on, name the authoritative source.0-30 daysChief data officer
Remove superseded documents from the index rather than leaving them to rank low.30-90 daysHead of data engineering
The evidence — 1 document
Organization
Document
Position
McKinsey & CompanyConsultancy · June 2026
Data readiness for scaling AI impactOur reading Proposes governance over the foundation rather than over each application, which is the level at which a conflict between sources can be settled at all.A governed and reusable foundation
proposes a fix
Who is represented
This dossier is drawn from 24 organizations working on the subject, 3 of which are cited directly in the issues above.
Consultancy — 7
McKinsey & Company 4Genpact 1Boston Consulting Group 2Deloitte 1Forrester 1QuantumBlack, AI by McKinsey 1UST 1
Institution — 3
NIST 3Cloud Security Alliance 2World Economic Forum 1