Skip to content

DATABRICKS

Read original ↗

Databricks — Evaluation-First AI Agents: How Zepto Scales Customer Support on Databricks and MLflow

Summary

Zepto — an Indian quick-commerce platform running customer support as a multi-agent AI system processing 100,000+ tickets/day — worked with Databricks to make evaluation the primary way agents are built, tested, and operated rather than shipping ever more agents. The thesis is that at scale the binding constraint shifts from raw model capability to system assurance: even a 1% error rate on 100K daily tickets produces thousands of bad outcomes. The framework is a dual-loop model — a development loop (design/iterate/regression-test candidate agent versions against a golden dataset) and a production loop (monitor live traces, score them, detect failures) — joined by a quality gate that only promotes a version that meets all pillar thresholds and beats the current production baseline. Underneath sits a decomposable agent architecture (an orchestrator/router over vertical intent-specialist agents and horizontal oversight agents), rich per-invocation tracing into Delta/Unity Catalog via OpenTelemetry spans, LLM-plus-rule scorers ("AI jury") calibrated against human labels, automated prompt optimization, model optionality, and risk-stratified sampling of production traces. Reported outcomes: 80%+ tickets AI-managed, 65% support cost reduction, 3× faster development, 4× faster resolution.

Key takeaways

  1. At scale the constraint is assurance, not capability. "Just ship the agent" breaks because agentic systems are multi-step workflows (classify → retrieve → analyze → reason → call tools → generate) where failures can emerge anywhere and the final answer hides internal errors. Traditional software monitoring is "completely useless for AI": an agent can have perfect uptime and zero errors while repeating the same wrong answer — "to the engineers, the dashboard looks green; to the customer, it's a disaster." (Source: this article.)
  2. Dual loop joined by a quality gate. A development loop (design/iterate/regression-test before shipping) and a production loop (monitor live behavior, detect failures) are connected by a quality gate that decides which versions enter production and which are pushed back — plus a feedback loop where production failures enrich the next iteration's dataset. (Source: this article — patterns/snapshot-replay-agent-evaluation.)
  3. Trace everything as the foundation (Phase 0). Every agent invocation emits a rich execution trace (prompts, completions, retrieved docs, tool calls, latencies, decision paths). Enabled with MLflow's mlflow.<library>.autolog() plus @mlflow.trace for custom spans; traces are emitted in real time as OpenTelemetry spans with unique IDs and centralized via MLflow's Unity Catalog integration into Delta tables (patterns/telemetry-to-lakehouse).
  4. Evaluation pillars turn multi-stakeholder debate into a measurable contract (Phase 1). Each stakeholder answers "what does success mean to you?"; the answers become pillars (customer experience, operational efficiency, risk & compliance, financial impact), each with numeric gates that must be met before deployment.
  5. The golden dataset is the cornerstone and compounds over time (Phase 2). The single source of truth for dev-loop evaluation covers normal, edge, and failure cases, richly annotated (scenario type, business line, risk level), with stakeholder expectations captured as examples (e.g. the security team contributes prompt-injection / identity-attack / data-exfiltration adversarial cases — see concepts/prompt-injection). Zepto's six-month trajectory: 500 examples / 8-point dev–prod accuracy gap → 2,000 / 2-point → 5,247 / 0.4-point. The dev–prod gap is itself a dataset-quality signal; every hour on dataset quality saves ~10 hours of production debugging (a "10× multiplier"), and every production failure adds a trace back into the golden dataset.
  6. Prompt engineering is automated, not hand-written (Phase 3). Using MLflow prompt optimization: register an initial prompt, generate and optimize variants against the same scorers that gate deployment, run A/B evals automatically, deploy the best. Critically, the optimizer reflects with a strong model but scores candidates with a cheaper one so the search stays cost-aware (patterns/cheap-approximator-with-expensive-fallback, patterns/prompt-optimizer-flywheel).
  7. Scorers as an "AI jury," calibrated against humans (Phase 4). Use LLM-based scorers only where human-like judgment is needed and deterministic rules where they suffice; calibrate judges against human labels to 80–90% agreement and use multiple judges for high-stakes decisions (patterns/human-calibrated-llm-labeling).
  8. Model optionality is a first-class component (Phase 5). The framework switches between proprietary and open-source models by changing a model name in Databricks model serving / the Foundation Model API, so the dev loop can continuously search the cost/performance/quality frontier (patterns/ai-gateway-provider-abstraction).
  9. Auto-regression makes promotion a repeatable gate (Phase 6). Any change to logic/prompt/model auto-triggers an evaluation with three inputs — the new version, the golden dataset, and the current production baseline. The gate checks both "meets all pillar thresholds?" and "at least as good as the baseline?"; pass → auto-promote, fail → reject and the existing agent keeps serving traffic.
  10. Risk-stratified sampling beats uniform sampling in the production loop (Phase 7). Evaluating 100% of traffic is too expensive, but naive 10% uniform sampling misses most edge cases. Sampling rates are weighted by high-value customers, new/recently-changed flows, negative sentiment / high escalation risk, and image-based / fraud-prone interactions — yielding an effective 18–20% sample (~14,400 traces/day) that captures 45–60% of edge cases and detects issues in 4–6 minutes. Reported vs uniform: 86% lower review cost per issue found, 9× better edge-case detection (patterns/stratified-evaluation-sampling).
  11. Decomposable agent architecture makes evaluation tractable. A query (chat or image) hits an orchestrator/router (which can hand off to a human at any point) over two agent classes: vertical agents — intent specialists (WIMO order tracking/ETA, Missing, Expiry, Returns, Quality, Unable-to-Pay, General fallback) — and horizontal agents — oversight layers that cut across use cases (Image Deduplication; Item Matching & Image-Manipulation Detection). Metrics compute per vertical (WIMO intent F1, Expiry OCR accuracy) and per horizontal (fraud precision, image-reuse/manipulation detection) (patterns/specialized-agent-decomposition).
  12. Never optimize a single metric. Chasing intent accuracy in isolation once gained 5 points of intent but raised latency 133% and cost 0.4 points of CSAT; a composite multi-objective score gained 3 points of intent at only 17% more latency and +0.2 CSAT. Multiple perspectives (LLM judges, rules, human labels, production metrics) "create truth together"; the redundancy is insurance.

Production stories (where evaluation earned its keep)

  • Story 1 — the ETA that never moved. An incident left riders stuck in traffic while the agent looped "arriving in 10 mins" because it read cached data. Token-counter + warning scorers caught the repetition and high escalation risk within 5 minutes, surfacing traces of stationary riders with unchanged ETAs → a rule change (if a rider is stationary >10 min, give an honest update and proactively offer cancellation for a full refund) → an entire "cancel on delay" feature line. One online-eval insight became a product feature.
  • Story 2 — cancellation regression caught pre-prod. Adding cancellation handling to the WIMO agent confused three intents ("where is my order" / "I want to cancel" / "was my order cancelled"). Dev-phase MLflow eval caught the drop (intent accuracy 92.1% → 87.4%, poor F1 on new WIMO_CANCEL / WIMO_CANCEL_STATUS) against the golden dataset — no customer saw it; prompt optimization + dataset updates restored 94.2% (better than baseline) and it shipped with zero rollbacks.
  • Story 3 — calibrating multimodal agents against human judgment. Produce-quality scoring is subjective (mushrooms rated 2/5 vs 3/5 by different raters). They measured rater disagreement with Cohen's Kappa and treated it as the reliability ceiling ("no model can be more consistent than the humans it learns from"). The AI "played it safe," piling scores at 3; instead of minimizing error against the mean they matched the shape of the human score distribution. Online eval also surfaced out-of-scope cases (curdled milk shelved as packaged goods; taste/smell complaints a photo can't show) routed to a separate path (concepts/multi-modal-attribute-extraction).
  • Story 4 — closing the abuse backdoors. Refund abuse used catalog images, edited photos, and images reused across claims. The multimodal pipeline ran preprocessing (blur/brightness/resolution) + OCR + a jury of three vision models with consensus rules to decide auto-approval vs human review, layered with blur/screenshot/duplicate detection, image-vs-SKU matching, image-vs-stated-reason checks, and proof-of-delivery validation (patterns/multimodal-content-understanding).

Operational numbers

  • Volume: 100,000+ AI-agent tickets/day; even 1% error = thousands of bad outcomes/day.
  • Outcomes: 80%+ tickets fully AI-managed with human oversight; 65% support-cost/ticket reduction; payback < 1 month; +20% CSAT; +8% accuracy; 3× faster dev cycles; 4× faster resolution.
  • Golden dataset growth: 500 → 2,000 → 5,247 examples; dev–prod accuracy gap 8 → 2 → 0.4 points over six months.
  • Production loop: stratified sampling 18–20% effective (~14,400 traces/day), 45–60% edge-case capture, issue detection 4–6 min; 86% lower review cost per issue, 9× better edge-case detection vs uniform sampling.
  • Alerting: critical alerts checked every 5 min (intent-accuracy drops, groundedness violations, high escalation risk, P95 latency breaches); high/medium alerts (empathy degradation, cost spikes, tool-failure rates, CSAT trends, fraud-detection rate, multimodal latency).
  • Feedback-loop automation saved ~155 hours/month (~two FTEs).
  • Judge calibration target: 80–90% agreement with human labels.
  • Single-metric trap: +5 intent alone → +133% latency, −0.4 CSAT; composite → +3 intent, +17% latency, +0.2 CSAT.

Caveats

  • First-party Databricks + Zepto post; all outcome numbers (65% cost, 4× resolution, 86%/9× sampling gains, $-equivalent FTE savings) are self-reported, not independently audited.
  • Tier-3 source and AI-agent/ML-methodology framed, but ingested because the durable content is production system-design: multi-agent orchestration, trace-to-lakehouse observability at 100K tickets/day, risk-stratified sampling, quality-gated promotion, and two concrete production-incident stories (one a classic stale-cache failure).
  • The architecture is described at the block-diagram level; the orchestrator/router internals, exact model roster, and scorer implementations are not disclosed.
  • "Evaluation-first" is presented as a general blueprint; the specific thresholds and pillar weights are Zepto-specific and not generalizable verbatim.

Extracted vocabulary

Source

Last updated · 766 distilled / 2,225 read