Data readiness
The 8th-largest theme in the research: 66 of 268 documents, 36 organizations. Decomposed into 5 topics, each with its own issues, evidence and steps.
66 documents in the theme
12 issues named
22 sourced citations
5 organizations cited
0 sourced statistics
The topics
Each opens a dossier with the same shape: what the state of it is, the issues ranked by how many independent organizations name them - or marked as our own analysis where none does - the disagreements, the numbers, and the concrete steps under each remedy.
Unstructured data & retrieval
Most of what an enterprise knows is not in a table, and getting at it is its own engineering problem.
A document is treated as one object, and its meaning is lost at the first step 1
The unstructured problem is solved and the answer still cannot be trusted 1
77 documents · 37 organizations · 2 issues · 5 steps
Lineage, provenance & the golden source
Where a number came from, and whether two systems agree on it.
In a chain of agents, a weak signal upstream does not average out 2
Lineage is recorded at the door and lost through the pipeline ours
Two retrieval paths give two answers and neither is wrong ours
35 documents · 23 organizations · 3 issues · 8 steps
What "AI-ready" actually means
The distance between the data an organization holds and data a model can use.
Two thirds of the data is not in a state a model can use 2
Up to two fifths of the working week already goes on fixing the data 1
Remediating the estate, or remediating the paths the work runs on ours
28 documents · 18 organizations · 3 issues · 8 steps
Ownership & data products
Who is accountable for a dataset, and whether it is run as a product or a side effect.
The dataset everything depends on is owned by nobody 1
Each team rebuilds the same pipeline, and each one differs ours
17 documents · 10 organizations · 2 issues · 6 steps
Rights, consent & permitted use
Whether you are allowed to use the data this way, and where it may sit.
Access to a dataset is not the same as permission to train on it ours
Whether a deletion request reaches what has already been embedded ours
15 documents · 12 organizations · 2 issues · 6 steps
Where the sources disagree
No contradictions recorded yet.
From the toolkit
Data readiness — the implementation kit
The whole of Data readiness, turned into something you can run. A diagnostic that tells you which of these problems you have, and an action plan with every step owned and time-boxed.
A diagnostic you can run in a room 12 questions across 5 areas, each written so a yes or no tells you whether you have that problem. No scoring model to learn.
An action plan that names who does it 33 actions, each carrying a role and a time-box, and every fix states what exists when it is done - so you can tell a fix that landed from one that was attempted.
The same actions, sorted by person An owner map, so one column goes to one person, and a sequence that says what to do first rather than leaving you to guess.
Written for your situation Three editions - listed company, private company or scale-up, and advisory - so the owner names match the room you are actually in.
Yours to use in front of a client Every word is original work. No third-party research is reproduced in it, which is what makes it safe to hand on.
12 diagnostic questions · 33 owned actions · 22 pages · one-off, updates included
Issues were named by hand after reading the documents cited under each one. Consensus counts distinct organizations, not documents, and counts only evidence a human has verified against a located passage.
Every link opens the publishing organization's own page. Summaries and characterisations are written here; no publisher prose is reproduced.