Ownership & data products

Who is accountable for a dataset, and whether it is run as a product or a side effect.

17documents on this topic
10organizations represented
2issues named
5sourced citations
0sourced statistics

The state of it

One of 5 topics within Data readiness.

Weak governance and ownership is named as a root cause of data debt, and it is the one that cannot be bought. Tooling gaps close with a purchase. Fragmentation closes with engineering. An unowned dataset stays unowned through both.

The failure has a recognisable shape: a dataset is produced as a side effect of a system that exists for another purpose, consumed by three teams who each patch it locally, and owned by nobody who can decide anything about it. When a model starts depending on it, the absence of an owner becomes the reason nothing gets fixed - there is no one to ask, and everyone who could fix it has a different local workaround that already works for them.

The corrective in the material is to run data as a product: a named owner, consumers who are known, a defined quality level, and a route to change it. The word matters less than the four properties, and the four properties are what an unowned dataset lacks.

The issues, by agreement

How many independent organizations name each issue as a problem. An issue is only as real as the number of separate publishers that identify it, so the count is the ranking. Bars are organizations, not documents. Where the count reads ours, no publisher here states the issue and the analysis is our own.

Where they disagree

No contradictions recorded on this topic yet.

The issues in full

Each issue carries the organizations that name it, the numbers behind it, and the remedies proposed - with the concrete steps under each. Every citation points at a section of a named document, so any count here can be checked.

Issue 011 organization name it2026 evidence

The dataset everything depends on is owned by nobody

A dataset is produced as a by-product of a system built for something else, consumed by several teams, patched locally by each of them, and owned by none of them.

Weak ownership is named as a root of data debt alongside fragmentation and legacy architecture, and it is the one that survives every technical fix. The diagnostic question is not who maintains the pipeline; it is who is allowed to decide that a field changes meaning, and who has to be told. If that person does not exist, each consumer builds a local correction, the corrections diverge, and a model trained or retrieving across them inherits every one of them at once.

How to fix it — 1 approach, 3 steps

Give every dataset a named owner with the authority to change it

Each dataset a funded use case depends on gets one named human owner, a known consumer list, a stated quality level and a route to change it.

Done when Every dataset a funded use case depends on names one accountable person rather than a team, with a stated quality level and a consumer list that is notified before a field changes meaning.

  1. List the datasets the funded use cases depend on and who currently maintains each.0-30 daysChief data officer
  2. Name one accountable owner per dataset, not a team, and tell the consumers.30-90 daysChief data officer
  3. Require consumers to be notified before a field changes meaning.90-180 daysHead of data engineering
The evidence — 3 documents
OrganizationDocumentPosition
GenpactConsultancy · June 2026$18 trillion in trapped AI valueOur reading Names weak governance and ownership among the structural roots of data debt, alongside fragmented sources, legacy architecture, poor quality management and tooling gaps.Weak governance and ownership as a root causenames it
Boston Consulting GroupConsultancy · November 2024Data & Digital Platforms playbookOur reading Places ownership in the operating model for the platform rather than leaving it to whichever team built the pipeline.Ownership inside the platform operating modelproposes a fix
McKinsey & CompanyConsultancy · June 2026Data readiness for scaling AI impactOur reading Proposes publishing curated data as products with defined access routes, which presumes somebody owns each one.Data published as productsproposes a fix

Issue 02Our analysis2026 evidence

Each team rebuilds the same pipeline, and each one differs

Ingestion, extraction, quality checking, metadata and indexing are rebuilt per project. The mechanics differ slightly each time, so quality and lineage differ too, and none of it compounds.

This is the practical reason path-by-path remediation is affordable in one organization and ruinous in another. Where a common extensible pattern exists, each new workflow reuses the mechanics and inherits the quality checks and the lineage. Where it does not, the second workflow costs what the first did, and the two produce differently-shaped data that nobody can reconcile later. The pattern is the asset, more than any individual pipeline built with it.

How to fix it — 1 approach, 3 steps

Build the pipeline pattern once and require its reuse

Make the ingestion-to-index mechanics a shared, extensible pattern, and require new work to use it rather than rebuild it.

Done when A shared ingestion-to-index pattern covering quality, metadata and lineage is documented and in use, and every pipeline either uses it or has a recorded exception with its reason.

  1. Compare two existing pipelines and list where their quality checks differ.0-30 daysHead of data engineering
  2. Extract one extensible pattern covering ingestion, quality, metadata and lineage.30-90 daysHead of data engineering
  3. Require new pipelines to use it, and record any exception and its reason.90-180 daysHead of architecture
The evidence — 2 documents
OrganizationDocumentPosition
Boston Consulting GroupConsultancy · November 2024Data & Digital Platforms playbookOur reading Proposes reusable components as the mechanism by which platform investment compounds instead of repeating.Reusable platform componentsproposes a fix
McKinsey & CompanyConsultancy · June 2026Data readiness for scaling AI impactOur reading Describes teams reusing one extensible pattern for ingestion, extraction, quality checks, metadata, lineage and indexing rather than each building its own.A common, extensible pipeline patternproposes a fix

Who is represented

This dossier is drawn from 11 organizations working on the subject, 3 of which are cited directly in the issues above.

Consultancy — 6

McKinsey & Company 6 Boston Consulting Group 2 Genpact 1 Accenture 1 L.E.K. Consulting 1 UST 1

Hyperscaler — 3

IBM 2 Google Cloud 1 ServiceNow 1

Enterprise — 1

Grab 1

Vendor — 1

IntuitionLabs 1