How we eliminated $1 million a year of wasted AI agent spend in one hour¶
Summary¶
Databricks engineers rely heavily on AI agents that call MCP tool servers (Jira, Google Drive/Docs, wikis, log/usage tables). When a tool call misbehaves, the calling agent rarely fails loudly — it retries, guesses, and works around the problem, silently burning tokens and developer wait time while the task still appears to complete. Using Unity Gateway's OpenTelemetry tracing (a single unified trace table over every MCP invocation) plus Genie One to query those traces in plain English, they found seven small bugs in their tool servers costing an estimated $499K/year in wasted tokens and ~12,000 engineering hours/year in agent wait time — ~$1.2M/year in lost productivity total. Finding, quantifying, and fixing all seven took about an hour. The deeper lesson: most of these were not the model "calling the tool wrong" — the tool accepted only one of several reasonable input shapes and crashed on the rest, so the fix is to build tools that adapt to how LLMs naturally call them.
Key takeaways¶
- Silent tool failures are a first-class, hidden cost center. From the outside the task still completes; on an aggregate dashboard the waste looks like a ~10% bump in token spend that is easily misread as usage growth (silent-tool-failure-cost). (Source: sources/2026-09-01-databricks-how-we-eliminated-1-million-a-year-of-wasted-ai-agent-spend)
- The gateway already had the data. Unity Gateway sits on the path of every call and automatically emits an OpenTelemetry trace per MCP tool invocation — tool name, arguments, error (if any), token counts, latency, and a session ID tying calls together — landing in one unified trace table. No new instrumentation was required.
- Genie One turned schema-spelunking + SQL into plain-English questions. Pointed at the trace table, it answered "which tool errors recur the most / how many turns to recover / what does each cost in tokens and wall-clock" in minutes. "Most of our hour went to reading those answers rather than writing queries." (trace-driven-tool-failure-diagnosis)
- Recovery cost tracks error-message quality almost perfectly. Self-documenting messages → ~14% repeat rate, ~4.6 turns to recover; cryptic tracebacks → ~30% repeat, ~12 turns; misleading messages → ~50% repeat, ~13 turns (error-message-quality-drives-agent-recovery-cost).
- The model usually wasn't wrong — the tool was too rigid. MCP tool signatures are kept deliberately under-specified (for generality and to save the per-call token cost of parameter descriptions). When a signature is vague, the model fills the gap with a reasonable guess (e.g. a JSON array for "a list of fields"), and a tool that accepts only one shape crashes on the rest (tool-selection-accuracy). Design principle: coerce the list into a string, default the omitted parameter, absorb the unexpected argument — honor the flexibility the loose signature promised (tools-adapt-to-how-llms-call-them).
- The fix was the cheap part. Once Genie One handed over a ranked list of what to fix and what the model was actually sending, applying the fixes across tool servers was a quick pass with a coding agent. "The scarce, expensive step was never writing the fix. It was knowing what to fix."
- The loop is repeatable: Unity Gateway makes agent behavior observable; Genie One makes it queryable without SQL. Trace the calls, ask what keeps going wrong, fix, repeat.
Operational numbers¶
Single 24-hour window (Jira + Google Drive/Docs tool servers). Annualized:
| Bug | Errors/day | Annual token cost | Annual wait time | Repeat rate |
|---|---|---|---|---|
Jira: KeyError: 'fields' (get) |
137 | $250K | 2,500 h | ~30% |
Jira: 'list' object has no attribute 'split' |
535 | $87K | 4,850 h | 30.5% |
Jira: KeyError: 'fields' (search) |
32 | $58K | 580 h | ~30% |
| GDrive: Invalid field selection | 417 | $46K | 2,740 h | 54.5% |
Jira: unexpected analysis_prompt kwarg |
121 | $42K | 840 h | 50.0% |
GDocs: find_text required |
137 | $15K | 440 h | 14.3% |
Jira: quote_from_bytes() expected bytes |
30 | $1.2K | 73 h | 66.7% |
| Total | 1,409 | $499K | 12,023 h | — |
- Total lost productivity ~$1.2M/year ($499K tokens + ~12,000 engineering hours of agent wait time).
- Highest-volume bug (535 failures/day):
issues.searchexpected a comma-separated string like"key,summary,status"and called.split(); the model naturally passed a JSON array (a list). A list has no.split()→ raw traceback'list' object has no attribute 'split'. Average 12 turns to recover, 30% of sessions hit it more than once → ~$87K/yr and 4,850 h from a single.split()call. - Google Drive Invalid field selection: 49.6% of all
drive_file_getcalls failed because the model passed valid-looking Drive API field names (id,name,mimeType) the tool's endpoint didn't accept.
Error-message quality vs recovery cost¶
| Error message quality | Example | Repeat rate | Avg turns to recover |
|---|---|---|---|
| Self-documenting | "find_text and replace_text required" | 14% | 4.6 |
| Somewhat informative | "Missing required parameters: org, repo" | ~30% | 4 |
| Cryptic traceback | "'list' object has no attribute 'split'" | 30.5% | 12.1 |
| Misleading | "unexpected keyword argument 'analysis_prompt'" | 50% | 13.1 |
Caveats¶
- First-party Databricks engineering blog; all dollar and hour figures are Databricks' own estimates extrapolated from a single 24-hour trace window, not an independently audited study.
- Token-cost and wait-time attributions are modeled from trace data (per-error token counts × recovery turns), not measured against a controlled baseline.
- The tool servers, models, and coding agent involved are Databricks-internal; the exact models and harnesses are unnamed.
Source¶
- Original: https://www.databricks.com/blog/how-we-eliminated-1-million-year-wasted-ai-agent-spend-one-hour
- Raw markdown:
raw/databricks/2026-09-01-how-we-eliminated-1-million-a-year-of-wasted-ai-agent-spend-2b139af1.md
Related¶
- systems/unity-ai-gateway — emits the per-tool-call OTel trace table this analysis ran on
- systems/databricks-genie — Genie One queried the trace table in natural language
- systems/opentelemetry — the trace standard the gateway emits
- systems/model-context-protocol — the tool-server protocol whose loose signatures set up the failures
- silent-tool-failure-cost
- error-message-quality-drives-agent-recovery-cost
- tool-selection-accuracy
- concepts/token-overhead
- trace-driven-tool-failure-diagnosis
- tools-adapt-to-how-llms-call-them