Skip to content

DATABRICKS

Read original ↗

How we eliminated $1 million a year of wasted AI agent spend in one hour

Summary

Databricks engineers rely heavily on AI agents that call MCP tool servers (Jira, Google Drive/Docs, wikis, log/usage tables). When a tool call misbehaves, the calling agent rarely fails loudly — it retries, guesses, and works around the problem, silently burning tokens and developer wait time while the task still appears to complete. Using Unity Gateway's OpenTelemetry tracing (a single unified trace table over every MCP invocation) plus Genie One to query those traces in plain English, they found seven small bugs in their tool servers costing an estimated $499K/year in wasted tokens and ~12,000 engineering hours/year in agent wait time — ~$1.2M/year in lost productivity total. Finding, quantifying, and fixing all seven took about an hour. The deeper lesson: most of these were not the model "calling the tool wrong" — the tool accepted only one of several reasonable input shapes and crashed on the rest, so the fix is to build tools that adapt to how LLMs naturally call them.

Key takeaways

  • Silent tool failures are a first-class, hidden cost center. From the outside the task still completes; on an aggregate dashboard the waste looks like a ~10% bump in token spend that is easily misread as usage growth (silent-tool-failure-cost). (Source: sources/2026-09-01-databricks-how-we-eliminated-1-million-a-year-of-wasted-ai-agent-spend)
  • The gateway already had the data. Unity Gateway sits on the path of every call and automatically emits an OpenTelemetry trace per MCP tool invocation — tool name, arguments, error (if any), token counts, latency, and a session ID tying calls together — landing in one unified trace table. No new instrumentation was required.
  • Genie One turned schema-spelunking + SQL into plain-English questions. Pointed at the trace table, it answered "which tool errors recur the most / how many turns to recover / what does each cost in tokens and wall-clock" in minutes. "Most of our hour went to reading those answers rather than writing queries." (trace-driven-tool-failure-diagnosis)
  • Recovery cost tracks error-message quality almost perfectly. Self-documenting messages → ~14% repeat rate, ~4.6 turns to recover; cryptic tracebacks → ~30% repeat, ~12 turns; misleading messages → ~50% repeat, ~13 turns (error-message-quality-drives-agent-recovery-cost).
  • The model usually wasn't wrong — the tool was too rigid. MCP tool signatures are kept deliberately under-specified (for generality and to save the per-call token cost of parameter descriptions). When a signature is vague, the model fills the gap with a reasonable guess (e.g. a JSON array for "a list of fields"), and a tool that accepts only one shape crashes on the rest (tool-selection-accuracy). Design principle: coerce the list into a string, default the omitted parameter, absorb the unexpected argument — honor the flexibility the loose signature promised (tools-adapt-to-how-llms-call-them).
  • The fix was the cheap part. Once Genie One handed over a ranked list of what to fix and what the model was actually sending, applying the fixes across tool servers was a quick pass with a coding agent. "The scarce, expensive step was never writing the fix. It was knowing what to fix."
  • The loop is repeatable: Unity Gateway makes agent behavior observable; Genie One makes it queryable without SQL. Trace the calls, ask what keeps going wrong, fix, repeat.

Operational numbers

Single 24-hour window (Jira + Google Drive/Docs tool servers). Annualized:

Bug Errors/day Annual token cost Annual wait time Repeat rate
Jira: KeyError: 'fields' (get) 137 $250K 2,500 h ~30%
Jira: 'list' object has no attribute 'split' 535 $87K 4,850 h 30.5%
Jira: KeyError: 'fields' (search) 32 $58K 580 h ~30%
GDrive: Invalid field selection 417 $46K 2,740 h 54.5%
Jira: unexpected analysis_prompt kwarg 121 $42K 840 h 50.0%
GDocs: find_text required 137 $15K 440 h 14.3%
Jira: quote_from_bytes() expected bytes 30 $1.2K 73 h 66.7%
Total 1,409 $499K 12,023 h —
  • Total lost productivity ~$1.2M/year ($499K tokens + ~12,000 engineering hours of agent wait time).
  • Highest-volume bug (535 failures/day): issues.search expected a comma-separated string like "key,summary,status" and called .split(); the model naturally passed a JSON array (a list). A list has no .split() → raw traceback 'list' object has no attribute 'split'. Average 12 turns to recover, 30% of sessions hit it more than once → ~$87K/yr and 4,850 h from a single .split() call.
  • Google Drive Invalid field selection: 49.6% of all drive_file_get calls failed because the model passed valid-looking Drive API field names (id, name, mimeType) the tool's endpoint didn't accept.

Error-message quality vs recovery cost

Error message quality Example Repeat rate Avg turns to recover
Self-documenting "find_text and replace_text required" 14% 4.6
Somewhat informative "Missing required parameters: org, repo" ~30% 4
Cryptic traceback "'list' object has no attribute 'split'" 30.5% 12.1
Misleading "unexpected keyword argument 'analysis_prompt'" 50% 13.1

Caveats

  • First-party Databricks engineering blog; all dollar and hour figures are Databricks' own estimates extrapolated from a single 24-hour trace window, not an independently audited study.
  • Token-cost and wait-time attributions are modeled from trace data (per-error token counts × recovery turns), not measured against a controlled baseline.
  • The tool servers, models, and coding agent involved are Databricks-internal; the exact models and harnesses are unnamed.

Source

Last updated · 766 distilled / 2,225 read