Unstructured data & retrieval

Most of what an enterprise knows is not in a table, and getting at it is its own engineering problem.

77documents on this topic
37organizations represented
2issues named
4sourced citations
0sourced statistics

The state of it

One of 5 topics within Data readiness.

Most of what an organization knows is not in a table. It is in contracts, tickets, call recordings, slide decks and PDFs, and the assumption that this is simply a matter of pointing a model at a folder is where a great deal of money goes.

The most useful engineering account here makes a point that sounds pedantic and is not: a source artifact stops being a single static object. One PDF becomes extracted text, tables, images and image summaries - several derived objects, each of which has to stay linked to the original and to each other if meaning, lineage and access control are to survive the trip. Treat the document as one blob and all three are lost at the first step, invisibly, and everything downstream inherits the loss.

The corrective is to give unstructured data the treatment structured data has had for decades - ingestion, extraction, quality checks, metadata, lineage, indexing - and to publish the result as curated products reachable by text, metadata, vector search and API, rather than as a folder somebody points a retriever at.

And it is explicit in the same source that solving unstructured data alone is not enough: the value comes from connecting it to the structured estate in one governed foundation, which is the harder half and the one usually deferred.

The issues, by agreement

How many independent organizations name each issue as a problem. An issue is only as real as the number of separate publishers that identify it, so the count is the ranking. Bars are organizations, not documents. Where the count reads ours, no publisher here states the issue and the analysis is our own.

Where they disagree

No contradictions recorded on this topic yet.

The issues in full

Each issue carries the organizations that name it, the numbers behind it, and the remedies proposed - with the concrete steps under each. Every citation points at a section of a named document, so any count here can be checked.

Issue 011 organization name it2026 evidence

A document is treated as one object, and its meaning is lost at the first step

A PDF goes into the pipeline as a file and comes out as text. The tables, the images, the structure and the relationship between them are gone, and nothing downstream can tell what was lost.

This is the failure that produces retrieval which is confidently wrong: the number was in a table, the table became prose, and the model answers from the prose. Handled properly one source artifact yields several derived objects - extracted text, tables, images, summaries of images - each retained, each linked back to the original and across to the others, so that meaning, lineage and control travel with the content instead of being stripped at ingestion. It is more work at the start and it is the difference between a retrieval system that can be audited and one that cannot.

How to fix it — 1 approach, 3 steps

Keep every derived object, and keep it linked to its source

Process each source artifact into its parts, retain them, and preserve the link back to the original and across the parts.

Done when Tables and images from source documents are retained as their own objects, and a retrieved passage comes back with a resolvable link to the original document.

  1. Take one document type and list what is lost between the source and the index today.0-30 daysHead of data engineering
  2. Retain tables and images as their own objects, linked to the source document.30-90 daysHead of data engineering
  3. Require the link to the original to be returned with every retrieved passage.90-180 daysHead of architecture
The evidence — 3 documents
OrganizationDocumentPosition
McKinsey & CompanyConsultancy · June 2026Data readiness for scaling AI impactOur reading Describes one PDF producing extracted text, tables, images and image summaries, each kept linked to the original and to one another so that meaning, lineage and control survive the pipeline.A source artifact is no longer a single objectnames it
McKinsey & CompanyConsultancy · June 2026Data readiness for scaling AI impactOur reading Proposes publishing the result as curated products reachable by text, metadata and vector search as well as APIs, so applications retrieve the right passage, table or entity before any model is called.Curated unstructured data productsproposes a fix
UberEnterpriseEnhanced Agentic-RAG: chatbots at near-human precisionOur reading Sets out the retrieval side as something engineered and measured rather than assumed to follow from having the documents.Retrieval quality as an engineering problemproposes a fix

Issue 021 organization name it2026 evidence

The unstructured problem is solved and the answer still cannot be trusted

Effort concentrates on documents because that is where the novelty is. The answers that matter need the documents joined to the structured record, and that join is left for later.

It is stated plainly in the source and it is the half that gets deferred: connecting structured and unstructured into one governed, traceable, reusable foundation is the requirement, not either alone. The reason it slips is that the unstructured work demos well and the join does not - a retrieval system over contracts looks impressive on its own, right up to the point where somebody asks which of those contracts are still active, and the answer lives in a system the retriever never touched.

How to fix it — 1 approach, 2 steps

Prove one joined answer before scaling the document pipeline

Require at least one production answer that depends on both the documents and the structured record before the retrieval work is scaled.

Done when One production answer is in service that could not be produced from the documents or the system of record alone, and the document set was not widened before it worked.

  1. Pick one question that needs both a document and a system of record to answer.0-30 daysHead of architecture
  2. Build that path end to end, including the join, before widening the document set.30-90 daysHead of data engineering
The evidence — 1 document
OrganizationDocumentPosition
McKinsey & CompanyConsultancy · June 2026Data readiness for scaling AI impactOur reading States that solving unstructured data on its own is insufficient, and that readiness requires structured and unstructured data connected into one governed, traceable and reusable foundation.Unstructured alone will not be enoughnames it

Who is represented

This dossier is drawn from 37 organizations working on the subject, 2 of which are cited directly in the issues above.

Consultancy — 11

McKinsey & Company 5 Accenture 4 Boston Consulting Group 4 Capgemini 4 Deloitte 4 EY 4 KPMG 2 Arthur D. Little 1 Booz Allen Hamilton 1 Infosys 1 UST 1

Institution — 6

NIST 2 World Economic Forum 2 Association of Corporate Counsel 1 Citi 1 Cloud Security Alliance 1 FinOps Foundation 1

Academic — 3

arXiv (research) 1 Carnegie Mellon SEI 1 National Bureau of Economic Research 1

Hyperscaler — 9

AWS 5 IBM 5 Google Cloud 4 Microsoft 3 Fujitsu 1 Google 1 Lenovo 1 OpenText 1 Samsung SDS 1

Frontier lab — 2

Anthropic 3 OpenAI 1

Enterprise — 3

Uber 3 Pinterest 1 Vimeo 1

Vendor — 2

Menlo Ventures 2 Writer 1

Other — 1

CISA 1