When it fails: response and recovery

What happens after something goes wrong - who is told, what is undone, and what changes as a result.

41documents on this topic
25organizations represented
3issues named
4sourced citations
0sourced statistics

The state of it

One of 6 topics within Evaluation & observability.

Four topics in this theme are about noticing that something is wrong. This one is about the half-hour after you notice, and it is the thinnest-covered ground in the whole index - which is itself the finding.

The response route that exists was built for attackers. Read across the research, the vocabulary of incident response is almost entirely cyber: 15 documents discuss security incidents in depth, and they are good - practised plans, named scenarios, regulatory reporting clocks, escalation to a security function that runs this for a living. That machinery is real and it works.

What is missing is the other failure. A system that was not attacked, was not breached, and was confidently wrong - at scale, to customers, for a fortnight before anyone noticed. Probing 291 documents for the language that failure would need, almost nothing comes back: redress or a right to contest appears in 1, customer harm in none, withdrawing or rolling back a deployed model in 1 - and that one is about credentials, not models. Escalation appears in 19, but escalation is the route in; it says who gets told, not what they can undo.

Two honest qualifications. This is a statement about the research held here, not about the world - serious-incident reporting duties and risk-management frameworks exist in regulation, and none of them is held here. And a gap in what publishers write about is not proof of a gap in what organizations do. But the tier-1 advisory material an executive would actually reach for is close to silent on it, and that is worth knowing before assuming the playbook covers you.

The practical consequence is narrow and testable: the security incident process will not fire, because nothing was breached. So the question to ask is not whether there is an incident plan. It is who has the authority to switch a working system off, how the people already given a wrong answer are found, and what obligates anything to change afterwards.

The issues, by agreement

How many independent organizations name each issue as a problem. An issue is only as real as the number of separate publishers that identify it, so the count is the ranking. Bars are organizations, not documents. Where the count reads ours, no publisher here states the issue and the analysis is our own.

Who takes which position

The chart above counts positions; this shows whose they are. Read down a column for what one organization holds across the whole topic, and across a row for who lines up on one issue. Where a cell carries more than one position, the strongest is shown and the rest are in the tooltip.

Ddisputes it Qqualifies it Nnames it as a problem Pproposes a fix
When it fails: response and recovery: 3 issues against the 4 organizations cited on them. The number under each name is how many of these issues it is cited on.
Issue Boston Consulting Group · 1 Carnegie Mellon SEI · 1 Cloud Security Alliance · 1 Microsoft · 1
The plan exists on paper and has never been run · P N P
The route back to people who were already given a wrong answer P · · ·
The failure route assumes an attacker, so being confidently wrong triggers nothing · · · ·

A dot means this organization is not cited on that issue. It does not mean they are silent on it: an organization is cited where its document takes a position we could locate, and the absence of a citation is the absence of a finding, not a finding of absence. Who is represented lists everyone working on this topic, including those not cited above.

Where they disagree

No contradictions recorded on this topic yet.

The issues in full

Each issue carries the organizations that name it, the numbers behind it, and the remedies proposed - with the concrete steps under each. Every citation points at a section of a named document, so any count here can be checked.

Issue 011 organization name it2026 evidence

The plan exists on paper and has never been run

A response plan is written, approved and filed. It has not been rehearsed against a named scenario, so the first execution is during a real event.

The security material is unambiguous that the plan has to be developed, tested and practised against specific scenarios, and it is the one place this discipline is stated plainly enough to borrow. Rehearsal is what finds the things a written plan cannot: that the person named has left, that the log does not go back far enough, that nobody agrees who can authorise switching it off on a Saturday. The same argument applies with more force to the non-adversarial case, because there is no security function that runs it monthly out of habit.

How to fix it — 1 approach, 3 steps

Rehearse one scenario per system per year, and fix what it breaks

Run the response once against a named scenario, with the actual people, and treat what fails in the rehearsal as the real output.

Done when Each live system has run one named failure scenario in the last twelve months with the named owners rather than deputies, timed step by step, and what the rehearsal broke has been fixed and re-run.

  1. Pick one named failure scenario per live system and schedule a rehearsal.0-30 daysHead of risk
  2. Run it with the named owners, not deputies, and time each step.30-90 daysChief operating officer
  3. Fix what the rehearsal broke, and re-run the parts that failed.90-180 daysHead of risk
The evidence — 3 documents
OrganizationDocumentPosition
Cloud Security AllianceInstitution · July 2026AI security through the CISO lensOur reading Names the distance between a documented capability and a demonstrated one.Preparedness in practicenames it
Carnegie Mellon SEIAcademicThe AI Adoption Maturity ModelOur reading Places post-incident learning inside an assessed capability rather than leaving it to goodwill, which is what makes it survive a quiet quarter.Learning as a capability that is assessedproposes a fix
MicrosoftHyperscalerDigital defense reportOur reading Proposes rehearsal against specific named scenarios rather than a written plan alone, and ties readiness to statutory reporting timeframes.Develop, test and practise the planproposes a fix

Issue 02Our analysisnewest evidence Nov 2024

The route back to people who were already given a wrong answer

Attention goes to fixing the system. Whether anything identifies who received bad output while it was broken, tells them, and corrects what followed from it is not addressed by the sources here.

This is the part with almost no published support anywhere in the research, and it is the part that turns a technical fault into a customer or regulatory problem. It needs three things that are unglamorous and have to exist beforehand: a log good enough to enumerate who was affected, a decision about the threshold at which they are told, and somebody empowered to reverse a decision the system already made. Retrofitting the first one after an incident is usually impossible, which is why it belongs in the design. No organization in this index names the missing route back. It is raised here because the incident material consistently stops at restoring the system, and the consensus count is zero for that reason.

How to fix it — 1 approach, 3 steps

Be able to list who was affected, and decide the telling threshold now

Ensure every deployed system can produce the list of who received its output over any window, and agree in advance what level of error obliges telling them.

Done when Each live system can produce the list of who received its output over any date range, the error level that obliges telling them is written down in advance, and the person who can reverse a decision is named.

  1. Confirm each live system can enumerate recipients of its output for any date range.0-30 daysHead of data
  2. Agree the threshold at which affected people are contacted, and write it down.30-90 daysHead of risk
  3. Name who can reverse a decision the system already made, and what that costs.30-90 daysChief operating officer
The evidence — 1 document
OrganizationDocumentPosition
Boston Consulting GroupConsultancy · November 2024Exec playbook on AI riskOur reading Proposes locating this kind of obligation in the risk and compliance function rather than in the delivery team, which is where it otherwise falls by default.Risk and compliance in the operating modelproposes a fix

Issue 03Our analysis

The failure route assumes an attacker, so being confidently wrong triggers nothing

Incident processes are inherited from security, where the trigger is a breach. A system that is simply wrong breaches nothing, so no threshold is crossed and no process starts.

This is a gap in coverage rather than a disagreement between publishers. The security material is strong and specific, and where AI appears in it, it appears as a new attack surface - prompt injection, tool invocation, data poisoning - which is a real problem and a different one. The non-adversarial failure has no owner by default: it is not a security event, it rarely produces an outage, and the people who notice first are usually the ones receiving the wrong output rather than the ones running the system.

How to fix it — 1 approach, 3 steps

Name who owns a failure that is not an attack

Write down, for each deployed system, who is accountable when it is wrong rather than breached, and what authority that person has.

Done when Each live system names an accountable owner for being wrong rather than breached, that owner can suspend it without further approval, and the trigger is stated as a number rather than a judgment.

  1. Name an accountable owner for non-security failure of each live system.0-30 daysChief operating officer
  2. Give that owner explicit authority to suspend the system without a further approval.30-90 daysCEO
  3. Define the trigger: what level of wrongness starts the process, stated as a number.30-90 daysHead of risk
The evidence — 0 documents
OrganizationDocumentPosition

Who is represented

This dossier is drawn from 25 organizations working on the subject, 4 of which are cited directly in the issues above.

Consultancy — 8

Boston Consulting Group 2 McKinsey & Company 2 Accenture 1 Deloitte 1 EY 1 Genpact 1 Infosys 1 PwC 1

Institution — 5

Cloud Security Alliance 4 NIST 3 World Economic Forum 3 Association of Corporate Counsel 1 FS-ISAC 1

Academic — 1

Carnegie Mellon SEI 1

Hyperscaler — 5

Microsoft 2 Google Cloud 3 IBM 3 AWS 2 OpenText 1

Frontier lab — 2

Anthropic 1 OpenAI 1

Vendor — 3

Palo Alto Networks 2 7AI 1 Predibase 1

Other — 1

CISA 1