Defining what good looks like

The quality bar, written down before the system is built rather than after.

13documents on this topic
11organizations represented
2issues named
7sourced citations
0sourced statistics

The state of it

One of 6 topics within Evaluation & observability.

13 documents from 11 organizations address defining what good looks like. This is the smallest practitioner-shaped topic in the research and the one whose absence causes the most downstream trouble.

The single most portable idea in the whole set sits here, and it comes from OpenAI: classify every result as ready to use, needing correction, or needing escalation, then divide the full cost - including human review, retries and rework - by the number that met the bar. It measures the thing a business actually feels, it is comparable across vendors and systems, and it can be adopted without any new tooling. Almost nothing else in the research offers a definition that portable.

The issues, by agreement

How many independent organizations name each issue as a problem. An issue is only as real as the number of separate publishers that identify it, so the count is the ranking. Bars are organizations, not documents. Where the count reads ours, no publisher here states the issue and the analysis is our own.

Where they disagree

No contradictions recorded on this topic yet.

The issues in full

Each issue carries the organizations that name it, the numbers behind it, and the remedies proposed - with the concrete steps under each. Every citation points at a section of a named document, so any count here can be checked.

Issue 011 organization name it2026 evidence

The quality bar is set after the demonstration rather than before it

Teams build, demonstrate, and then decide what counts as good enough - which guarantees the bar sits just below what was achieved.

Setting the standard after seeing the result is not dishonesty, it is the default behaviour of any project under pressure to show progress. The correction is procedural and cheap: write the acceptance criterion before work starts, in the investment paper, and have it approved by someone who is not delivering it. OpenAI's guidance is explicit that the quality bar should be defined before testing begins.

How to fix it — 2 approaches, 6 steps

Write the acceptance criterion into the funding paper

No build starts without a written definition of the standard and who signs off that it was met.

Done when The AI investment template carries a mandatory acceptance criterion, each is approved by someone outside the delivery team before work starts, and any change to a criterion is recorded against the original.

  1. Add a mandatory acceptance-criterion field to the AI investment template.0-30 daysCFO
  2. Have it approved by someone outside the delivery team before work begins.0-30 daysRisk
  3. Measure against the original criterion, and restate history if it is ever changed.ongoingCFO

Adopt one portable outcome scale

Ready to use, needs correction, needs escalation. One scale across every system makes results comparable and vendor claims checkable.

Done when Every AI system reports on the same three-outcome scale, cost is reported per accepted outcome rather than per call, and vendor evaluations arrive on that scale.

  1. Adopt the three-outcome classification as the reporting standard for all AI systems.0-30 daysCIO
  2. Report cost per accepted outcome rather than cost per call or per token.30-90 daysCFO
  3. Require vendors to report against the same scale in evaluations.90-180 daysProcurement
The evidence — 4 documents
OrganizationDocumentPosition
OpenAIFrontier lab · July 2026How to control AI usage and spendOur reading Puts the sequencing in writing as a rule: what counts as an acceptable answer is settled before any testing begins, and spend at the validation stage is released only against a standard that already exists.Evaluate model efficiency by outcome ROInames it
AnthropicFrontier lab · December 2024Building effective agentsOur reading Proposes using separate model calls to evaluate different aspects of performance rather than having the acting call judge its own output.Automating evalsproposes a fix
Google CloudHyperscalerGoogle’s agentic AI transformation frameworkOur reading Proposes formal pilot gate reviews that approve progression against success criteria, which puts the bar at the gate rather than after the demonstration.MVP/Pilot gate reviews against success criteriaproposes a fix
OpenAIFrontier lab · August 2026A scorecard for the AI ageOur reading Proposes classifying every result as ready to use, needing correction, or needing escalation, and dividing full cost - including employee time, human review, retries and rework - by the number of tasks that met the bar.Three outcomes; cost per successful taskproposes a fix

Issue 021 organization name it1 qualifies it2026 evidence

The model was evaluated and the system is what shipped

Selection effort concentrates on choosing a model. What reaches a customer is a model wrapped in retrieval, prompts, tools and orchestration - and almost none of that is what was measured.

Microsoft, reporting from a forum of 250 customers running AI at scale, describes the shift plainly: in the first wave the model was the decision and organizations treated choosing it as the work, until it became clear that the harness around it mattered as much - the data it can reach, the context it is handed, the infrastructure it runs on. Every one of those is a place the system can be wrong while the model is right. A benchmark score is a statement about a component; the failure a customer meets is a property of the assembly. This is why a system can pass every evaluation it was given and still be unusable, and why swapping in a better model so often changes nothing.

How to fix it — 1 approach, 3 steps

Evaluate the assembly, on your own cases

Score the whole path a request takes, using cases from your business. A model benchmark tells you nothing about the system you built around it.

Done when A test set built from real requests already through the business process runs against the full system rather than the model, with a record of where it fails, and it re-runs whenever any part of the harness changes.

  1. Build a test set from real requests that have already been through your process.0-30 daysCTO
  2. Run it against the full system, not the model, and record where it fails.30-90 daysCTO
  3. Re-run the same set whenever any part of the harness changes, not only the model.ongoingCTO
The evidence — 3 documents
OrganizationDocumentPosition
MicrosoftHyperscaler · June 2026Tokenomics is the new headcountOur reading Records the change of view among leaders running this at scale: the model was treated as the decision until the surrounding harness proved to matter as much, and the number of parts required to deliver value keeps growing.The system matters more than the modelnames it
AnthropicFrontier labBuilding multi-agent researchOur reading Qualifies it with a measurement: most of the difference in outcome came from how much was spent traversing the problem and which tools were called, not from the choice of model alone.What explained the performance variancequalifies it
AWSHyperscaler · October 2024Best practices for building gen AI agentsOur reading Places the design work in the assembly - tools, memory, control flow - which is the part a model benchmark never covers.Designing the agent, not just selecting the modelproposes a fix

Who is represented

This dossier is drawn from 12 organizations working on the subject, 5 of which are cited directly in the issues above.

Consultancy — 1

McKinsey & Company 1

Institution — 3

Moody’s 1 NIST 1 World Economic Forum 1

Academic — 1

Carnegie Mellon SEI 1

Hyperscaler — 4

Google Cloud 2 Microsoft 2 AWS 1 IBM 2

Frontier lab — 2

Anthropic 3 OpenAI 2

Enterprise — 1

LinkedIn 1