Skip to content

CONCEPT Cited by 1 source

Token overhead

Definition

Token overhead is the portion of an AI coding agent's per-inference cost that comes from context the user did not explicitly type — the tool outputs, codebase searches, retrieved files, skills, and system information the harness injects around the user's actual request.

When a user types something short ("Please investigate and fix this bug"), the agent then gathers massive amounts of context, invokes many tools, searches the codebase, and integrates company-provided skills. By the time the (costly) LLM inference runs, the user's original statement is a negligible fraction of the input — costs are dominated by the surrounding overhead. (Source: sources/2026-08-07-databricks-managing-ai-coding-costs-at-scale)

Why it matters

  • Token overhead, not user input, is the dominant cost term in agentic coding. Optimizing it is a distinct lever from switching to a cheaper model (concepts/efficiency-frontier) or routing (patterns/ai-gateway-provider-abstraction).
  • The techniques are described as "still new," but even simple tuning yields large wins: at Databricks, tuning harness + caching settings produced an "almost 50% reduction in the number of generated tokens and associated costs, with no observed quality degradation."

Techniques for reducing it

  • Coerce more frequent compaction (compression) of the active context (context-compaction-for-serving-cost).
  • Use less-chatty harnesses or tune existing harnesses to emit less token overhead.
  • Audit popular tools and reduce their verbosity — tool outputs are a major overhead source.
  • Break tasks into smaller units of work to shrink context scope per task.
  • Tune prompt caching. Cache writes cost money but cached reads drastically cut per-inference cost; hand-tuning cache settings to raise hit rate is workload-dependent and can be very impactful (prompt-cache-tuning-for-cost).

Relationship to context engineering

Token overhead is the cost view of the same phenomenon that context engineering treats from a quality view: the assembled context is the real driver of both agent behavior and agent cost. Reducing overhead (compaction, verbosity audits, smaller tasks) overlaps heavily with context-engineering discipline — the context window is a scarce, managed resource.

Seen in

Last updated · 766 distilled / 2,225 read