Evaluation & observability
The fourth-largest theme in the research: 173 of 268 documents, 62 organizations. Decomposed into 6 topics, each with its own issues, evidence and steps.
173 documents in the theme
13 issues named
34 sourced citations
18 organizations cited
4 sourced statistics
The topics
Each opens a dossier with the same shape: what the state of it is, the issues ranked by how many independent organizations name them - or marked as our own analysis where none does - the disagreements, the numbers, and the concrete steps under each remedy.
When it fails: response and recovery
What happens after something goes wrong - who is told, what is undone, and what changes as a result.
The plan exists on paper and has never been run 1
The failure route assumes an attacker, so being confidently wrong triggers nothing ours
The route back to people who were already given a wrong answer ours
41 documents · 25 organizations · 3 issues · 9 steps
Monitoring in production
Whether anyone is watching the system after it goes live.
Every team chose its own stack, so there is nothing to observe across them 2
Monitoring watches the infrastructure and not the behaviour 1
There is a dashboard, and no threshold that triggers anything 1
and 1 more
38 documents · 21 organizations · 4 issues · 14 steps
Testing before deployment
What is checked before it reaches a customer, and by whom.
Testing is a launch gate for a system that keeps changing after launch 1
21 documents · 20 organizations · 1 issue · 3 steps
Trust & independent verification
Who checks the claim, and whether they are independent of the team making it.
The team that built it is the only team able to verify it 1
16 documents · 14 organizations · 1 issue · 3 steps
Defining what good looks like
The quality bar, written down before the system is built rather than after.
The model was evaluated and the system is what shipped 1
The quality bar is set after the demonstration rather than before it 1
13 documents · 11 organizations · 2 issues · 9 steps
Choosing the metric
Which number gets reported, and whether it tracks anything the business feels.
The reported number is the one that is easy to collect 2
One reliability bar is set for the whole estate, when the bar is a function of what fails 1
7 documents · 7 organizations · 2 issues · 8 steps
Where the sources disagree
No contradictions recorded yet.
From the toolkit
Evaluation & observability — the implementation kit
The whole of Evaluation & observability, turned into something you can run. A diagnostic that tells you which of these problems you have, and an action plan with every step owned and time-boxed.
A diagnostic you can run in a room 13 questions across 6 areas, each written so a yes or no tells you whether you have that problem. No scoring model to learn.
An action plan that names who does it 46 actions, each carrying a role and a time-box, and every fix states what exists when it is done - so you can tell a fix that landed from one that was attempted.
The same actions, sorted by person An owner map, so one column goes to one person, and a sequence that says what to do first rather than leaving you to guess.
Written for your situation Three editions - listed company, private company or scale-up, and advisory - so the owner names match the room you are actually in.
Yours to use in front of a client Every word is original work. No third-party research is reproduced in it, which is what makes it safe to hand on.
13 diagnostic questions · 46 owned actions · 25 pages · one-off, updates included
Issues were named by hand after reading the documents cited under each one. Consensus counts distinct organizations, not documents, and counts only evidence a human has verified against a located passage.
Every link opens the publishing organization's own page. Summaries and characterisations are written here; no publisher prose is reproduced.