What is checked before it reaches a customer, and by whom.
21documents on this topic
20organizations represented
1issues named
4sourced citations
0sourced statistics
The state of it
One of 6 topics within Evaluation & observability.
21 documents from 20 organizations address what is checked before a system reaches a customer. The material is dominated by pre-deployment validation and red teaming, and it carries an assumption that does not hold for these systems: that testing establishes a property which then persists.
It does not. The model can be updated by a vendor, the prompt can be edited, the retrieved content changes daily, and the system is non-deterministic to begin with. A test suite run once at launch describes a system that no longer exists a month later. Almost nothing here addresses re-testing on a cadence, which is the control that would actually hold.
The issues, by agreement
How many independent organizations name each issue as a problem. An issue is only as real as the number of separate publishers that identify it, so the count is the ranking. Bars are organizations, not documents. Where the count reads ours, no publisher here states the issue and the analysis is our own.
Each issue carries the organizations that name it, the numbers behind it, and the remedies proposed - with the concrete steps under each. Every citation points at a section of a named document, so any count here can be checked.
Issue 011 organization name itnewest evidence Jul 2026
Testing is a launch gate for a system that keeps changing after launch
Pre-deployment validation is treated as establishing a durable property. The model version, the prompt and the retrieved content all change afterwards, and the suite is rarely run again.
Re-run the suite on every change, including the vendor's
Model version, prompt and retrieved content are all inputs. A change to any of them should re-trigger the tests automatically.
Done when Model version, prompt and retrieved content are each versioned inputs that trigger the suite automatically, the vendor is contractually obliged to give notice of a model change, and release is blocked on regression against the previous run.
Treat model version, prompt and retrieved content as versioned inputs that trigger the suite.0-30 daysCTO
Require the vendor to give notice before a model change reaches production.30-90 daysProcurement
Block release on regression against the previous run, not against an absolute bar.30-90 daysCTO
The evidence — 4 documents
Organization
Document
Position
Booz Allen HamiltonConsultancy · December 2024
Countering enterprise AI security threatsOur reading Describes adversaries targeting different phases of the AI system lifecycle with methods designed to degrade, deny, deceive or manipulate, which is a continuing exposure rather than a launch-time one.Attacks across the AI system lifecycle
names it
OpenAIFrontier lab · July 2026
How to control AI usage and spendOur reading Proposes defining the quality bar before testing and running evaluations that reflect real tasks including edge cases, then measuring the full cost of reaching that standard - a recurring measurement rather than a gate.Evaluate model efficiency by outcome
proposes a fix
UberEnterprise
Uber - Journey to Generative AIOur reading Describes tooling that measures model quality across all stages rather than at a single approval point.Quality measured across all stages
proposes a fix
UC BerkeleyAcademic · May 2025
An AI governance maturity matrix for boardsOur reading Proposes periodically auditing models for fairness with documented data sources and decision logic, explicitly on a recurring basis.Dimension 4: periodic auditing
proposes a fix
Who is represented
This dossier is drawn from 22 organizations working on the subject, 4 of which are cited directly in the issues above.
Consultancy — 5
Booz Allen Hamilton 1Capgemini 1Deloitte 1EY 1McKinsey & Company 1