Trust & independent verification

Who checks the claim, and whether they are independent of the team making it.

16documents on this topic
14organizations represented
1issues named
4sourced citations
1sourced statistics

The state of it

One of 6 topics within Evaluation & observability.

16 documents from 14 organizations address who verifies the claim. The structural answer is well understood and rarely implemented: the function that built the system should not be the only function attesting that it works.

For AI this is harder than for finance, and the research is honest about why. Verifying a model performance claim requires the skills to reproduce an evaluation, and those skills currently sit almost entirely inside the team being verified. Accenture's figures put the gap in context: only 39% say governance responsibilities are effectively embedded, and while 65% have integrated real-time data into decision-making, only 53% apply that timely data in governance forums.

The issues, by agreement

How many independent organizations name each issue as a problem. An issue is only as real as the number of separate publishers that identify it, so the count is the ranking. Bars are organizations, not documents. Where the count reads ours, no publisher here states the issue and the analysis is our own.

names it as a problemdisputes itqualifies it

Where they disagree

No contradictions recorded on this topic yet.

The issues in full

Each issue carries the organizations that name it, the numbers behind it, and the remedies proposed - with the concrete steps under each. Every citation points at a section of a named document, so any count here can be checked.

Issue 011 organization name itnewest evidence Jun 2026

The team that built it is the only team able to verify it

Independent assurance requires someone outside the delivery team who can reproduce an evaluation result. In most organizations that person does not exist, so verification is self-reporting with extra steps.

96%Respondents expecting risk and compliance roles to expand as AI embedsMoody’s · Sep 2025
How to fix it — 1 approach, 3 steps

Test the capability by reproducing one number

Pick a performance figure that reached the board and have the second line reproduce it without help. The result is your assurance capability, measured.

Done when The second line has attempted, unaided, to reproduce one reported AI performance claim from the last quarter, and the result is written down along with what is being funded to close the gap it revealed.

  1. Select one reported AI performance claim from the last quarter.0-30 daysAudit Committee
  2. Have the second line attempt independent reproduction, unaided.30-90 daysRisk
  3. Fund the gap the exercise reveals rather than accepting self-reporting.90-180 daysCFO
The evidence — 4 documents
OrganizationDocumentPosition
McKinsey & CompanyConsultancy · June 2026Data readiness for scaling AI impactOur reading Names the retrieval problem directly: an answer assembled from fragments of many documents cannot be traced back to any one of them.Traceability of unstructured retrievalnames it
AnthropicFrontier lab · December 2024Building effective agentsOur reading Proposes handling guardrail evaluation in a separate model call from the one producing the response, on the grounds that it performs better than asking one call to do both - the same separation principle applied technically.Sectioning: separate the guardrail from the responseproposes a fix
DeloitteConsultancyStrategic governance of AIOur reading Positions independent assurance over AI as an emerging oversight responsibility for audit committees rather than a delivery-team function.Independent assurance as an audit committee responsibilityproposes a fix
NISTInstitution · January 2023AI Risk Management Framework 1.0Our reading Proposes that regular assessment involve internal experts who were not front-line developers on the system, or assessors outside it entirely, with people beyond the building team consulted in proportion to the risk being carried.MEASURE 1.3: who performs the assessmentproposes a fix

Who is represented

This dossier is drawn from 16 organizations working on the subject, 4 of which are cited directly in the issues above.

Consultancy — 5

Deloitte 1 McKinsey & Company 1 Accenture 1 Capgemini 1 UST 1

Institution — 5

NIST 2 Cloud Security Alliance 2 Association of Corporate Counsel 1 Australian Institute of Company Directors 1 World Economic Forum 1

Hyperscaler — 4

Google Cloud 2 AWS 1 IBM 1 Microsoft 1

Frontier lab — 1

Anthropic 2

Other — 1

CISA 1