The quality bar, written down before the system is built rather than after.
13documents on this topic
11organizations represented
2issues named
7sourced citations
0sourced statistics
The state of it
One of 6 topics within Evaluation & observability.
13 documents from 11 organizations address defining what good looks like. This is the smallest practitioner-shaped topic in the research and the one whose absence causes the most downstream trouble.
The single most portable idea in the whole set sits here, and it comes from OpenAI: classify every result as ready to use, needing correction, or needing escalation, then divide the full cost - including human review, retries and rework - by the number that met the bar. It measures the thing a business actually feels, it is comparable across vendors and systems, and it can be adopted without any new tooling. Almost nothing else in the research offers a definition that portable.
The issues, by agreement
How many independent organizations name each issue as a problem. An issue is only as real as the number of separate publishers that identify it, so the count is the ranking. Bars are organizations, not documents. Where the count reads ours, no publisher here states the issue and the analysis is our own.
Each issue carries the organizations that name it, the numbers behind it, and the remedies proposed - with the concrete steps under each. Every citation points at a section of a named document, so any count here can be checked.
Issue 011 organization name it2026 evidence
The quality bar is set after the demonstration rather than before it
Teams build, demonstrate, and then decide what counts as good enough - which guarantees the bar sits just below what was achieved.
Setting the standard after seeing the result is not dishonesty, it is the default behaviour of any project under pressure to show progress. The correction is procedural and cheap: write the acceptance criterion before work starts, in the investment paper, and have it approved by someone who is not delivering it. OpenAI's guidance is explicit that the quality bar should be defined before testing begins.
Write the acceptance criterion into the funding paper
No build starts without a written definition of the standard and who signs off that it was met.
Done when The AI investment template carries a mandatory acceptance criterion, each is approved by someone outside the delivery team before work starts, and any change to a criterion is recorded against the original.
Add a mandatory acceptance-criterion field to the AI investment template.0-30 daysCFO
Have it approved by someone outside the delivery team before work begins.0-30 daysRisk
Measure against the original criterion, and restate history if it is ever changed.ongoingCFO
Adopt one portable outcome scale
Ready to use, needs correction, needs escalation. One scale across every system makes results comparable and vendor claims checkable.
Done when Every AI system reports on the same three-outcome scale, cost is reported per accepted outcome rather than per call, and vendor evaluations arrive on that scale.
Adopt the three-outcome classification as the reporting standard for all AI systems.0-30 daysCIO
Report cost per accepted outcome rather than cost per call or per token.30-90 daysCFO
Require vendors to report against the same scale in evaluations.90-180 daysProcurement
The evidence — 4 documents
Organization
Document
Position
OpenAIFrontier lab · July 2026
How to control AI usage and spendOur reading Puts the sequencing in writing as a rule: what counts as an acceptable answer is settled before any testing begins, and spend at the validation stage is released only against a standard that already exists.Evaluate model efficiency by outcome ROI
names it
AnthropicFrontier lab · December 2024
Building effective agentsOur reading Proposes using separate model calls to evaluate different aspects of performance rather than having the acting call judge its own output.Automating evals
proposes a fix
Google CloudHyperscaler
Google’s agentic AI transformation frameworkOur reading Proposes formal pilot gate reviews that approve progression against success criteria, which puts the bar at the gate rather than after the demonstration.MVP/Pilot gate reviews against success criteria
proposes a fix
OpenAIFrontier lab · August 2026
A scorecard for the AI ageOur reading Proposes classifying every result as ready to use, needing correction, or needing escalation, and dividing full cost - including employee time, human review, retries and rework - by the number of tasks that met the bar.Three outcomes; cost per successful task
proposes a fix
Issue 021 organization name it1 qualifies it2026 evidence
The model was evaluated and the system is what shipped
Selection effort concentrates on choosing a model. What reaches a customer is a model wrapped in retrieval, prompts, tools and orchestration - and almost none of that is what was measured.
Microsoft, reporting from a forum of 250 customers running AI at scale, describes the shift plainly: in the first wave the model was the decision and organizations treated choosing it as the work, until it became clear that the harness around it mattered as much - the data it can reach, the context it is handed, the infrastructure it runs on. Every one of those is a place the system can be wrong while the model is right. A benchmark score is a statement about a component; the failure a customer meets is a property of the assembly. This is why a system can pass every evaluation it was given and still be unusable, and why swapping in a better model so often changes nothing.
Score the whole path a request takes, using cases from your business. A model benchmark tells you nothing about the system you built around it.
Done when A test set built from real requests already through the business process runs against the full system rather than the model, with a record of where it fails, and it re-runs whenever any part of the harness changes.
Build a test set from real requests that have already been through your process.0-30 daysCTO
Run it against the full system, not the model, and record where it fails.30-90 daysCTO
Re-run the same set whenever any part of the harness changes, not only the model.ongoingCTO
The evidence — 3 documents
Organization
Document
Position
MicrosoftHyperscaler · June 2026
Tokenomics is the new headcountOur reading Records the change of view among leaders running this at scale: the model was treated as the decision until the surrounding harness proved to matter as much, and the number of parts required to deliver value keeps growing.The system matters more than the model
names it
AnthropicFrontier lab
Building multi-agent researchOur reading Qualifies it with a measurement: most of the difference in outcome came from how much was spent traversing the problem and which tools were called, not from the choice of model alone.What explained the performance variance
qualifies it
AWSHyperscaler · October 2024
Best practices for building gen AI agentsOur reading Places the design work in the assembly - tools, memory, control flow - which is the part a model benchmark never covers.Designing the agent, not just selecting the model
proposes a fix
Who is represented
This dossier is drawn from 12 organizations working on the subject, 5 of which are cited directly in the issues above.