CONCEPT Cited by 1 source
Token overhead¶
Definition¶
Token overhead is the portion of an AI coding agent's per-inference cost that comes from context the user did not explicitly type — the tool outputs, codebase searches, retrieved files, skills, and system information the harness injects around the user's actual request.
When a user types something short ("Please investigate and fix this bug"), the agent then gathers massive amounts of context, invokes many tools, searches the codebase, and integrates company-provided skills. By the time the (costly) LLM inference runs, the user's original statement is a negligible fraction of the input — costs are dominated by the surrounding overhead. (Source: sources/2026-08-07-databricks-managing-ai-coding-costs-at-scale)
Why it matters¶
- Token overhead, not user input, is the dominant cost term in agentic coding. Optimizing it is a distinct lever from switching to a cheaper model (concepts/efficiency-frontier) or routing (patterns/ai-gateway-provider-abstraction).
- The techniques are described as "still new," but even simple tuning yields large wins: at Databricks, tuning harness + caching settings produced an "almost 50% reduction in the number of generated tokens and associated costs, with no observed quality degradation."
Techniques for reducing it¶
- Coerce more frequent compaction (compression) of the active context (context-compaction-for-serving-cost).
- Use less-chatty harnesses or tune existing harnesses to emit less token overhead.
- Audit popular tools and reduce their verbosity — tool outputs are a major overhead source.
- Break tasks into smaller units of work to shrink context scope per task.
- Tune prompt caching. Cache writes cost money but cached reads drastically cut per-inference cost; hand-tuning cache settings to raise hit rate is workload-dependent and can be very impactful (prompt-cache-tuning-for-cost).
Relationship to context engineering¶
Token overhead is the cost view of the same phenomenon that context engineering treats from a quality view: the assembled context is the real driver of both agent behavior and agent cost. Reducing overhead (compaction, verbosity audits, smaller tasks) overlaps heavily with context-engineering discipline — the context window is a scarce, managed resource.
Related¶
- concepts/context-engineering — deliberate assembly/pruning of agent context
- concepts/context-window-as-token-budget — the context window as a scarce managed resource
- concepts/efficiency-frontier — the complementary "cheaper model" cost lever
- prompt-cache-tuning-for-cost — tune cache TTL/settings to raise hit rate
- context-compaction-for-serving-cost — compaction to bound serving cost
Seen in¶
- sources/2026-08-07-databricks-managing-ai-coding-costs-at-scale — names token overhead as the dominant cost term and reports ~50% token reduction from harness/cache tuning.
- sources/2026-09-01-databricks-how-we-eliminated-1-million-a-year-of-wasted-ai-agent-spend — a species of token overhead from failed tool calls: silent retries around seven broken MCP tools added an estimated $499K/year in wasted tokens (silent-tool-failure-cost).
- sources/2026-09-02-github-how-we-make-ai-coding-more-cost-efficient — GitHub Copilot attacks token overhead across four harness levers (selective output compaction, unused-formatting removal, prompt compression, batched background completions), but insists the objective is the whole task not the tool call (local-metric-trap).
- sources/2026-09-03-spotify-portal-by-spotify-cut-my-claude-code-token-usage-by-90 — attacks token overhead at the source by routing the I/O half of the work off the frontier model entirely: a hook blocks large reads and delegates them to a cheap worker mode (io-vs-reasoning-work-split, delegate-io-to-cheaper-worker-model); ~90% bulk-read saving.