Context & memory

The finite budget every agent runs against, and what happens when it fills.

16documents on this topic
13organizations represented
2issues named
8sourced citations
0sourced statistics

The state of it

One of 6 topics within Agent orchestration.

16 documents from 13 organizations address the context an agent runs against. It is the smallest topic in this theme and the one closest to pure engineering, which is probably why the executive material barely touches it - and why the cost surprises land where they do.

Context is a finite budget spent on every turn. A long-running agent accumulates history until the budget is consumed, at which point behaviour changes rather than stops. Anthropic's multi-agent design is in large part an answer to this: subagents hold their own context windows and return compressed findings, which is a context management strategy before it is an intelligence strategy.

The issues, by agreement

How many independent organizations name each issue as a problem. An issue is only as real as the number of separate publishers that identify it, so the count is the ranking. Bars are organizations, not documents. Where the count reads ours, no publisher here states the issue and the analysis is our own.

Where they disagree

No contradictions recorded on this topic yet.

The issues in full

Each issue carries the organizations that name it, the numbers behind it, and the remedies proposed - with the concrete steps under each. Every citation points at a section of a named document, so any count here can be checked.

Issue 012 organizations name it2026 evidence

Context is a finite budget, and long-running agents spend it without anyone watching

Every turn re-sends accumulated history. Cost grows with conversation length rather than with work done, and behaviour degrades as the window fills rather than failing cleanly.

How to fix it — 2 approaches, 5 steps

Set a context and turn budget per run

Cap the turns and the context an agent may consume in a single run, and make exceeding it a handled outcome rather than a surprise.

Done when Every production agent has explicit per-run limits on turns, context size and tool calls, a defined behaviour at the limit, and an alert on runs that hit the ceiling with the rate tracked over time.

  1. Set explicit per-run limits on turns, context size and tool calls for every production agent.0-30 daysCTO
  2. Define what happens at the limit: summarise and continue, escalate, or stop.0-30 daysCTO
  3. Alert on runs that hit the ceiling, and treat a rising rate as a design signal.30-90 daysCIO

Compress on purpose rather than by truncation

Decide what survives a long run. Left alone, the system will drop whatever happens to be oldest, which is not the same as what matters least.

Done when What must persist across a long-running task is defined and stored outside the context window, and summarisation happens at named checkpoints rather than by whatever truncation drops first.

  1. Define what must persist across a long-running task and store it outside the context window.30-90 daysCTO
  2. Summarise deliberately at checkpoints instead of relying on truncation.30-90 daysCTO
The evidence — 4 documents
OrganizationDocumentPosition
AnthropicFrontier labBuilding multi-agent researchOur reading Describes subagents operating in parallel with their own context windows and condensing the most important tokens back to a lead agent, which is a context management strategy as much as a capability one.Subagents and context compressionnames it
InfosysConsultancy · June 2026AI token economics for executivesOur reading Frames architectural decisions - how calls are routed and how retrieval is governed - rather than unit price as the determinant of spend.Token economics for executivesnames it
DeloitteConsultancyNavigating AI spend dynamicsOur reading Proposes context window limits and API usage caps alongside budget alerts as the practical control on runaway consumption.Guardrails: context window limits and usage capsproposes a fix
UberEnterpriseUber - Journey to Generative AIOur reading Proposes codifying the operational disciplines into the platform so they apply uniformly rather than per project.Platform-level practiceproposes a fix

Issue 022 organizations name it2026 evidence

Cost compounds invisibly because one request is not one request

One request may be a two-line reply, a document read end to end, or an agent working for an hour. They differ by orders of magnitude, and no conventional forecast distinguishes them.

How to fix it — 1 approach, 3 steps

Forecast by workload shape, not by request count

Segment demand into chat, retrieval-heavy, and long-running agentic work, and forecast each separately. A blended per-request number will be wrong in both directions.

Done when Chat, retrieval-heavy and long-running agentic work each have a measured cost from at least one month of actuals, each is forecast separately with variance reported by shape, and long-running work has its own ceiling.

  1. Classify workloads into the three shapes and measure actual cost per shape for one month.0-30 daysCIO
  2. Forecast each shape separately and report variance by shape, not in aggregate.30-90 daysCFO
  3. Set separate ceilings for long-running agentic work, which is where the tail sits.30-90 daysCFO
The evidence — 4 documents
OrganizationDocumentPosition
Boston Consulting GroupConsultancy · July 2026How to manage AI token costsOur reading Describes cost per outcome compounding sharply and invisibly through forces traditional software and infrastructure forecasts miss: breadth and depth of adoption as users move from chat to multistep research and agent deployment, and task intensity, where a short answer and a long-running agent session both look like one request.Four forces behind compounding costnames it
InfosysConsultancy · June 2026AI token economics for executivesOur reading Places routing and retrieval governance rather than unit price as the determinants of spend.Token economics for executivesnames it
AirbnbEnterprise · October 2024Automation Platform evolution for genAIOur reading Describes building the platform-level controls that make long-running generative workloads operable rather than handling each workload individually.Automation platform evolutionproposes a fix
DeloitteConsultancyNavigating AI spend dynamicsOur reading Proposes context window limits and usage caps as the direct control on the intensity dimension.Guardrailsproposes a fix

Who is represented

This dossier is drawn from 15 organizations working on the subject, 6 of which are cited directly in the issues above.

Consultancy — 4

Deloitte 2 Boston Consulting Group 1 Infosys 1 McKinsey & Company 2

Institution — 2

FinOps Foundation 1 World Economic Forum 1

Hyperscaler — 2

AWS 1 Google Cloud 1

Frontier lab — 1

Anthropic 2

Enterprise — 4

Airbnb 1 Uber 1 Pinterest 1 Vimeo 1

Vendor — 1

Predibase 1

Other — 1

CISA 1