Who checks the claim, and whether they are independent of the team making it.
16documents on this topic
14organizations represented
1issues named
4sourced citations
1sourced statistics
The state of it
One of 6 topics within Evaluation & observability.
16 documents from 14 organizations address who verifies the claim. The structural answer is well understood and rarely implemented: the function that built the system should not be the only function attesting that it works.
For AI this is harder than for finance, and the research is honest about why. Verifying a model performance claim requires the skills to reproduce an evaluation, and those skills currently sit almost entirely inside the team being verified. Accenture's figures put the gap in context: only 39% say governance responsibilities are effectively embedded, and while 65% have integrated real-time data into decision-making, only 53% apply that timely data in governance forums.
The issues, by agreement
How many independent organizations name each issue as a problem. An issue is only as real as the number of separate publishers that identify it, so the count is the ranking. Bars are organizations, not documents. Where the count reads ours, no publisher here states the issue and the analysis is our own.
Each issue carries the organizations that name it, the numbers behind it, and the remedies proposed - with the concrete steps under each. Every citation points at a section of a named document, so any count here can be checked.
Issue 011 organization name itnewest evidence Jun 2026
The team that built it is the only team able to verify it
Independent assurance requires someone outside the delivery team who can reproduce an evaluation result. In most organizations that person does not exist, so verification is self-reporting with extra steps.
96%Respondents expecting risk and compliance roles to expand as AI embedsMoody’s · Sep 2025
How to fix it — 1 approach, 3 steps
Test the capability by reproducing one number
Pick a performance figure that reached the board and have the second line reproduce it without help. The result is your assurance capability, measured.
Done when The second line has attempted, unaided, to reproduce one reported AI performance claim from the last quarter, and the result is written down along with what is being funded to close the gap it revealed.
Select one reported AI performance claim from the last quarter.0-30 daysAudit Committee
Have the second line attempt independent reproduction, unaided.30-90 daysRisk
Fund the gap the exercise reveals rather than accepting self-reporting.90-180 daysCFO
The evidence — 4 documents
Organization
Document
Position
McKinsey & CompanyConsultancy · June 2026
Data readiness for scaling AI impactOur reading Names the retrieval problem directly: an answer assembled from fragments of many documents cannot be traced back to any one of them.Traceability of unstructured retrieval
names it
AnthropicFrontier lab · December 2024
Building effective agentsOur reading Proposes handling guardrail evaluation in a separate model call from the one producing the response, on the grounds that it performs better than asking one call to do both - the same separation principle applied technically.Sectioning: separate the guardrail from the response
proposes a fix
DeloitteConsultancy
Strategic governance of AIOur reading Positions independent assurance over AI as an emerging oversight responsibility for audit committees rather than a delivery-team function.Independent assurance as an audit committee responsibility
proposes a fix
NISTInstitution · January 2023
AI Risk Management Framework 1.0Our reading Proposes that regular assessment involve internal experts who were not front-line developers on the system, or assessors outside it entirely, with people beyond the building team consulted in proportion to the risk being carried.MEASURE 1.3: who performs the assessment
proposes a fix
Who is represented
This dossier is drawn from 16 organizations working on the subject, 4 of which are cited directly in the issues above.
Consultancy — 5
Deloitte 1McKinsey & Company 1Accenture 1Capgemini 1UST 1
Institution — 5
NIST 2Cloud Security Alliance 2Association of Corporate Counsel 1Australian Institute of Company Directors 1World Economic Forum 1