Synthesis · top quartile of 269 publisher resources
What the evidence says executives should actually do about AI
Consultancies, hyperscalers and frontier labs publish AI playbooks faster than anyone can read them. Nobody has read them all. Read across 269 of them, 76% published in 2025–26, the disagreements are smaller than the noise suggests — as is the list of things that demonstrably work.
Answers come from this index's own analysis — the named issues,
the sourced positions and the owned steps. Every line links to the publisher it
came from. Nothing here is generated from publishers' text.
00 — What changed
What’s new
Everything added to the index appears here automatically as it is added — this is a record, not a selection. Listed from Aug 13.
Share of the 269 sources that substantively address each theme, and how many distinct publishers raise it. Measured across every document, not hand-picked.
Theme prevalence across 269 sources
How many of the 269 sources substantively discuss each theme. This measures the index rather than the industry: a short bar can mean the theme is under-sourced here, not that publishers are quiet about it.
03 — Moving the needle
What is measurably working
Findings with a number attached and a named source. Every one of these is a structural change, not a tool purchase.
Rebuilding a workflow end to end
>50%
In the early-adopter teams on one bank’s 400-component core modernization, time and effort fell by more than half, with people moved to supervising agent squads that document, write, review and integrate code. Reported by the consultancy from its own engagement.
Among 100 sales leaders whose organizations already run agentic AI, 89% report a positive effect on growth, 87% on rep productivity and 61% on lead conversion. The gains cluster in the earliest commercial steps — identifying which accounts to chase and in what order — while nurture further down scores mid-range. Adopters only, and self-reported.
35% of companies said they had an AI strategy in place, and 78% of those reported returns from generative AI (Google Cloud, 2024). Separately, organizations with someone owning AI at C-suite level report 10% higher return on AI spend — IBM counted equivalent roles as well as the formal title, so this is not evidence for creating a new seat.
The pattern is consistent: effort spent adding AI to the existing shape of the work, or waiting for the economics to fix themselves.
Copilots bolted onto processes nobody redesigned first
Grafting AI onto task flows that should first have been eliminated or consolidated makes low-value work faster and leaves the cost structure where it was. BCG rates the companies furthest along on AI maturity as reporting roughly three times the cost reduction of laggards — self-reported, and measured only inside the areas where AI is applied.
Total spend is set by how the system is built, not by the list price. Ungoverned retrieval and a default to frontier models are the two decisions that dominate the bill, and both are configuration rather than rebuild. The unit of management has changed.
Rolling the copilot wider instead of rebuilding the work
General-purpose assistants deploy fast and cheaply, and the time they save is real — but it lands a few minutes at a time across thousands of people, where no P&L line records it. The changes that show up in results are built into a single function’s process, and McKinsey estimates most of those are still short of production.
Managing token spend the way you manage cloud spend
FinOps exposes infrastructure cost well and is the right starting point — but a token bill moves with prompt length, retrieved context, model choice and how long an agent keeps going before it settles. The same spend belongs in three places at once: capital where it builds reusable capability, operating where it runs internal work, cost of goods where it sits inside the product.
The four debts, and the failure modes nobody budgets for
Genpact and HFS put $18 trillion of recoverable value behind four interlocking debts. 85% of leaders say those debts actively limit AI value; over half have no funded plan to address them.
Data debt
33% of data is AI-ready
The gap between the data firms hold and what AI needs. Data quality failures delay, degrade or sink 42% of analytics and AI initiatives. Cited as the single biggest blocker.
Process debt
about 40% of the work week is manual
Workflows that are largely undocumented and hard to change — fewer than half are formally documented. Point an agent at one unaltered and it carries out the same flawed sequence, only faster and at scale.
Tech debt
42% of development-team time absorbed
Core enterprise systems run to about ten years old on average, and roughly two fifths of development-team time goes to maintaining the debt they carry rather than building anything new.
Talent debt
32% of workforce is AI-ready
The readiness gap that quietly compounds all three other debts — and the one whose cost never lands in a ledger of its own.
61% say their AI has already been compromised
Among 1,000 C-level executives, 61% say something in their AI estate — a model, its data, or the assets around it — was breached in the past year. In a separate survey of 300 security and IT professionals, only about a quarter are confident their organization can run a security strategy for AI; 21% say plainly that they are not, and half sit on the fence.
A trust failure has cost a fifth to a majority of market value
The range most often quoted for market capitalization lost after a serious breach of trust. Follow it back and it is older and broader than it looks: Deloitte relays it here from its own 2021 work on institutional trust, which in turn attributes it to earlier research. It predates generative AI and is not about AI at all — carried here because it is the order of magnitude an AI governance failure is being weighed against, not because anyone has measured that.
As the stack consolidates, dependence shifts from the platform to the reasoning layer that shapes how the organization thinks. Where the boundary sits around the company’s own IP, essential data and business rules is a chief-executive question, and not one to settle inside IT alone.
Token spend behaves nothing like a software licence. It scales with usage, it has no natural budget owner, and the crossover points that decide build-versus-rent arrive faster than most planning cycles.
Three-year total cost of ownership at scale
Deloitte’s modelled comparison at equivalent configuration and token volume: once demand is sustained, an owned AI factory is roughly 2.1× more cost-effective than metered API hosting over three years.
Metered API hostingOwned AI factory
API hosting
index 100
AI factory
index 48
The curve moves fast in both directions: Deloitte models a greater than 90% fall in cost per billion tokens between year one and year three for the owned path, against roughly 150% annual TCO growth as volume climbs. The decision is therefore about sustained demand, not today’s per-token price. Two architectural choices — how model calls are routed and how retrieval is governed — account for 78% of achievable savings, and programmes that put governance in before scaling spend 40–50% less.
07 — Where to focus
Seven moves, and who owns each
Ordered by how much the research agrees on them and how early they gate everything downstream.
01
CEO / CFO
Re-baseline the budget on 1:3:5
McKinsey finds the transformations that work put three dollars into reshaping how the work runs, and five into teaching people to run it, for every one spent on the technology. Read as a budget test, a plan without that shape is not funded.
McKinsey
02
CFO / CIO
Instrument token spend at workflow level before scaling
Meter tokens per workflow, not per platform, before scaling. FinOps was built for infrastructure and does not price a loop an agent runs; the unit that matters is cost per useful outcome, with an owner on every material workflow.
BCG · Infosys
03
CEO / COO
Rebuild a few workflows end to end instead of assisting twenty
Both publishers argue the value sits in eliminating and reordering steps rather than adding assistants to them, and that spreading AI thinly is what produces diffuse benefit and no earnings impact. Pick where value is concentrated and rebuild the whole path.
BCG · McKinsey
04
CAIO / CISO
Bound each agent before you widen it
Anthropic’s security team screens every agentic use case on four axes: which of its inputs an attacker could write, which actions it can take and under whose identity, how far the damage reaches if it goes wrong, and whether its actions can be told apart from a person’s in your logs. Give it the least reach that still completes the work.
Anthropic
05
CDO
Pay down data debt along the chosen workflows
McKinsey finds readiness is the link between structured and unstructured data rather than a cleanup of either alone, and that governance has to reach the embedding and retrieval layer instead of stopping at the document. Genpact ranks data debt alongside process debt and finds the two are created by the same manual work. Sequencing the paydown behind the workflows you are rebuilding is our own recommendation.
McKinsey · Genpact
06
CEO / Board
Close the board gap in person
Frame what AI means for value creation yourself, run hands-on sessions, and consider a transformation committee of the more fluent directors rather than waiting for the full board to catch up.
BCG
07
CTO
Preserve model liquidity
A modular, layered stack with a security perimeter around proprietary knowledge lets you adopt better models as they arrive. Decide sovereignty deliberately rather than inheriting it from a vendor roadmap.
Palantir · BCG
08 — Featured reading
Twenty-four worth reading in full
Ranked by publisher standing, how many of the consensus themes each document substantively covers, and depth of treatment — capped at two per organization so no single firm dominates. Summaries are ours, written after reading each document in full.
McKinsey & CompanyData, Cloud & Architecture · June 2026
States that one answer can be assembled from fragments of many documents at once, so the trail is not straightforward, and identifies unstructured data entering and exiting AI systems with little traceability as a major challenge.
Access to a dataset is not the same as permission to train on it
A document is treated as one object, and its meaning is lost at the first step
Audit trails are built to satisfy a regulator, not to investigate an incident
Anthropic's deputy CISO on the framework his team applies before approving an agentic use case: what untrusted content it ingests, what actions it can take and on whose behalf, what the blast radius is if it misbehaves, and whether its activity is observable at all.
Four questions make agentic risk assessable without becoming a blocking gate
Introduces an identity spectrum for deciding on whose authority an agent acts
Argues the opposite of consolidation: train people inside the delivery and data science teams to run their own reviews and to recognise when to escalate, on the reasoning that a single central function becomes the gate every review waits at, and that each hand-off across an organizational boundary costs something. Offers its own internal programme as the worked example, and is explicit that the posture has to be enablement rather than refusal.
Agents run on one-time authentication and inherited permissions that cross boundaries
An agent can be turned without any injection at all
Cyber risk is ranked in the top three and delegated out of the leadership team
Reports 36% higher AI ROI for Chief AI Officers running hub-and-spoke or centralized operating models than for those managing decentralized ones, and roughly twice as many pilots reaching production, from a survey of more than 600 such officers fielded in the first quarter of 2025.
No one person owns policy, infrastructure, outcome and risk together
One person is asked to be both advocate and guardian
Ownership is a coalition, and one seat of it is often hostile
Names the pattern precisely: someone with low privilege induces a high-privilege agent to do what they could not do themselves, and because the action genuinely executes under the agent's trusted identity, the record it leaves reads as legitimate and pushes detection further out. Adds the credential half - keys and tokens that are static, shared between agents, or weakly held let an attacker act as the agent rather than merely through it, which defeats behavioural guardrails and the monitoring tuned to normal behaviour at the same time.
Agents run on one-time authentication and inherited permissions that cross boundaries
Autonomy is being set once, as a binary, instead of earned incrementally
Nobody can say whose authority an agent acted under
Carnegie Mellon SEITalent, Workforce & Change · PDF
A 63-page assessment model built by Carnegie Mellon's Software Engineering Institute with Accenture. Five maturity levels are scored across eight dimensions covering strategy, workforce, workflow and risk on one side, and data, engineering, operations and ecosystem on the other. Piloted with Fortune 500 early adopters.
Measures discipline and repeatability rather than how much AI has been deployed
Each capability area carries explicit goals, practices and artefacts, so claims can be evidenced
OpenAIAI Economics & Cost Management · August 2026
Proposes a concrete measure that separates working systems from demonstrations: count only tasks meeting the quality bar, divide full cost - including employee time, human review, retries and rework - by that number, and track results as ready to use, needs correction, or needs escalation.
Adoption is measured because the vendor console reports it, not because it answers anything
Models degrade quietly, and the first signal is usually a customer
Non-determinism breaks the testing process the organization already has
Widens the picture by listing what a regulated firm is already tracking, and the length of that list undercuts the idea of three blocs - though the items are not equivalent regimes. In the United States alone it names a federal executive order, a federal bill, the NIST framework, a securities regulator notice, statutes in Colorado and Texas, and a privacy regulator's automated decision-making rules in California. The United Kingdom appears not as a statute but as a government department and two regulators publishing separately. Asia-Pacific carries national laws in South Korea and Japan, Australian guidance alongside a voluntary safety standard, a Singapore consultation aimed at financial institutions, two Hong Kong publications and an ASEAN guide. The list is offered as non-exhaustive and omits China entirely, so it is a picture of the tracking burden rather than a census.
Executives are held accountable for AI without an agreed standard of care
Models degrade quietly, and the first signal is usually a customer
Regulatory monitoring is nobody's standing job
GenpactMarket & Adoption Research · June 2026 · PDF
A survey of more than 2,000 executives across 16 industries, conducted with HFS Research, which quantifies four interlocking obstacles: data, process, technology and talent debt. Its contribution is sizing them individually and modelling what resolving them is worth.
Sizes each debt separately; data and process dominate the total
Models remediation as a growth programme rather than an efficiency exercise
Describes investing in production capabilities including safe model deployment and automatic model retraining, and tooling that measures model quality across all stages, precisely because manual maintenance does not scale.
Context is a finite budget, and long-running agents spend it without anyone watching
Models degrade quietly, and the first signal is usually a customer
Monitoring watches the infrastructure and not the behaviour
McKinsey & CompanyAgentic AI & Agent Engineering · August 2026
Records the reasoning directly: if the agent gives the wrong answer and the employee acts on it, it is still the employee's mistake, so people are reluctant to rely on tools they do not fully understand where accuracy, professional judgment and reputation are on the line.
A centre of excellence is created without deciding what it decides
Executives sponsor the programme without visibly changing how they work
The AI plan and the people plan are written by different people, at different speeds
Boston Consulting GroupCase Studies & Deployments · December 2024 · PDF
Reports leaders investing in few high-priority opportunities with roughly double the ROI impact instead of hundreds of use cases, backed by around twice the investment in AI capability and twice the people.
A faster task is not a faster process
Spreading investment across many use cases halves the return
The base rate is cancellation, and cost is the leading cause
Records that both sides found the activity independently and began containing it - while itself treating that detection as insufficient, since the monitoring in place during internal testing is named among the things to strengthen and the systems had already gained network access, escalated privilege and reached a third party's production database before anyone noticed.
Detection works; the reach available before it fires is the problem
The containment boundary is a single component nobody treats as load-bearing
The control is understood, documented, and switched off where the work happens
Describes subagents operating in parallel with their own context windows and condensing the most important tokens back to a lead agent, which is a context management strategy as much as a capability one.
Context is a finite budget, and long-running agents spend it without anyone watching
Multi-agent is being adopted as a maturity stage rather than chosen for a problem
Multi-agent systems work largely by spending more tokens, and the budget rarely says so
Boston Consulting GroupAI Economics & Cost Management · July 2026 · PDF
A short executive perspective on why AI cost programmes underdeliver. BCG reports only a small minority generating value at scale, with leaders achieving roughly three times the cost reduction of laggards, and identifies the traps: fragmented initiatives, and AI layered onto processes instead of removing steps.
Insists targets be set on P&L outcomes rather than productivity proxies
Pairs AI with conventional cost levers instead of treating it as a substitute
IBM's treatment of agent governance in regulated environments, comparing the regulatory trajectories emerging in Australia and the European Union rather than assuming a single global regime. Its most useful contribution is a test for deployments that are permitted but still ill-advised.
Compares divergent regulatory paths instead of assuming convergence
Introduces a lawful-but-unwise test that sits above pure compliance
Proposes an identity and session model in which a revocation is notified once and every enforcement point consults shared state, so that a terminated agent is blocked across protocols and integrations rather than per system.
Agents run on one-time authentication and inherited permissions that cross boundaries
Detection works; the reach available before it fires is the problem
Nobody can say whose authority an agent acted under
USTAgentic AI & Agent Engineering · November 2025 · PDF
Sets out a specific task-level set - completion rate, work reduction, cost and cycle time per task, safety violations - explicitly in place of technical model scores, which are equally unrelated to whether work got done.
Adoption is measured because the vendor console reports it, not because it answers anything
An indicator is named without a calculation method, a source, a frequency or an owner
Evaluation is budgeted as overhead and turns out to be most of the system
A three-layer framework covering vision and strategy, an agentic development lifecycle, and the horizontal capabilities underneath. Its distinguishing feature is operational: each stage names objectives, inputs, activities, owners and deliverables, so it can be run as a programme rather than read as a point of view.
Every stage specifies roles and outputs, which makes it auditable
Governance, security and data access are treated as scaling prerequisites, not later additions
Notes that non-human identities are created by systems rather than business processes and run continuously without direct human oversight: a person reading a screen and pausing between actions creates natural moments to catch a mistake, while an identity executing thousands of operations a second offers no such window.
Agents take actions, and the chain leading to them cannot be reconstructed
Autonomy widens the gap between the decision and the person answerable for it
Risk to people rises with autonomy, and the trade is almost never stated
UC BerkeleyBoard Oversight & Governance · May 2025 · PDF
Names shortage of AI knowledge among board members as one of the most common hindrances to robust AI governance; oversight becomes cursory and directors defer to management.
Activity is being mistaken for a change in the operating model
AI reaches the board when something happens, not on a fixed cadence
Directors rate their own AI literacy as adequate; the executives reporting to them do not agree
GartnerBuild-Buy-Borrow & Vendor Strategy · March 2025
Qualifies the case for concentration by holding that AI now arrives from three directions at once - embedded in software already owned, bought by individual business units, and built centrally - so the centre's job is to govern the mix rather than to own all of it.
AI arrives through three doors, and the operating model watches one
Fear of falling behind is making the sourcing decision