Testing before deployment

What is checked before it reaches a customer, and by whom.

21documents on this topic
20organizations represented
1issues named
4sourced citations
0sourced statistics

The state of it

One of 6 topics within Evaluation & observability.

21 documents from 20 organizations address what is checked before a system reaches a customer. The material is dominated by pre-deployment validation and red teaming, and it carries an assumption that does not hold for these systems: that testing establishes a property which then persists.

It does not. The model can be updated by a vendor, the prompt can be edited, the retrieved content changes daily, and the system is non-deterministic to begin with. A test suite run once at launch describes a system that no longer exists a month later. Almost nothing here addresses re-testing on a cadence, which is the control that would actually hold.

The issues, by agreement

How many independent organizations name each issue as a problem. An issue is only as real as the number of separate publishers that identify it, so the count is the ranking. Bars are organizations, not documents. Where the count reads ours, no publisher here states the issue and the analysis is our own.

names it as a problemdisputes itqualifies it

Where they disagree

No contradictions recorded on this topic yet.

The issues in full

Each issue carries the organizations that name it, the numbers behind it, and the remedies proposed - with the concrete steps under each. Every citation points at a section of a named document, so any count here can be checked.

Issue 011 organization name itnewest evidence Jul 2026

Testing is a launch gate for a system that keeps changing after launch

Pre-deployment validation is treated as establishing a durable property. The model version, the prompt and the retrieved content all change afterwards, and the suite is rarely run again.

How to fix it — 1 approach, 3 steps

Re-run the suite on every change, including the vendor's

Model version, prompt and retrieved content are all inputs. A change to any of them should re-trigger the tests automatically.

Done when Model version, prompt and retrieved content are each versioned inputs that trigger the suite automatically, the vendor is contractually obliged to give notice of a model change, and release is blocked on regression against the previous run.

  1. Treat model version, prompt and retrieved content as versioned inputs that trigger the suite.0-30 daysCTO
  2. Require the vendor to give notice before a model change reaches production.30-90 daysProcurement
  3. Block release on regression against the previous run, not against an absolute bar.30-90 daysCTO
The evidence — 4 documents
OrganizationDocumentPosition
Booz Allen HamiltonConsultancy · December 2024Countering enterprise AI security threatsOur reading Describes adversaries targeting different phases of the AI system lifecycle with methods designed to degrade, deny, deceive or manipulate, which is a continuing exposure rather than a launch-time one.Attacks across the AI system lifecyclenames it
OpenAIFrontier lab · July 2026How to control AI usage and spendOur reading Proposes defining the quality bar before testing and running evaluations that reflect real tasks including edge cases, then measuring the full cost of reaching that standard - a recurring measurement rather than a gate.Evaluate model efficiency by outcomeproposes a fix
UberEnterpriseUber - Journey to Generative AIOur reading Describes tooling that measures model quality across all stages rather than at a single approval point.Quality measured across all stagesproposes a fix
UC BerkeleyAcademic · May 2025An AI governance maturity matrix for boardsOur reading Proposes periodically auditing models for fairness with documented data sources and decision logic, explicitly on a recurring basis.Dimension 4: periodic auditingproposes a fix

Who is represented

This dossier is drawn from 22 organizations working on the subject, 4 of which are cited directly in the issues above.

Consultancy — 5

Booz Allen Hamilton 1 Capgemini 1 Deloitte 1 EY 1 McKinsey & Company 1

Institution — 4

NIST 2 Citi 1 IAB 1 World Economic Forum 1

Academic — 2

UC Berkeley 1 Carnegie Mellon SEI 1

Hyperscaler — 5

AWS 1 Fujitsu 1 Google Cloud 1 IBM 1 Microsoft 1

Frontier lab — 2

OpenAI 2 Anthropic 1

Enterprise — 2

Uber 1 Grab 1

Vendor — 1

Palo Alto Networks 1

Other — 1

CISA 1