PATTERN Cited by 2 sources
Stratified evaluation sampling¶
Stratified evaluation sampling is the production-loop pattern of scoring only a fraction of live traffic, with per-stratum sampling rates weighted by risk and value instead of a single uniform rate — so that a fixed evaluation budget is concentrated on the interactions most likely to be wrong, costly, or dangerous.
Intent¶
Scoring 100% of production traffic with LLM judges and rule checks is prohibitively expensive at scale. The naive alternative — uniform random sampling at some low rate (e.g. 10%) — is cheap but blind: because failures and abuse concentrate in a small subset of traffic, uniform sampling spends most of its budget re-confirming the easy, healthy majority and misses most edge cases.
Stratification fixes the allocation: partition traffic into strata by risk/value and sample each stratum at a rate proportional to how much you need to see it. The total sample size stays within budget, but its composition shifts toward the interactions that matter.
Mechanism¶
- Instrument every interaction as a trace (the substrate — see patterns/telemetry-to-lakehouse / patterns/snapshot-replay-agent-evaluation).
- Define strata by risk and value. At Zepto the sampling rate is raised for: high-value customers; new or recently-changed flows; negative sentiment / high escalation risk; and image-based or fraud-prone interactions (Source: sources/2026-09-09-databricks-evaluation-first-ai-agents-how-zepto-scales-customer-support).
- Assign per-stratum sampling rates so the effective overall rate stays within the cost envelope while high-risk strata approach full coverage.
- Score the sampled traces with the same scorers that gate deployment (LLM judges + deterministic rules), write results to a governed store, and drive dashboards + tiered alerts (alert rules).
- Feed detected failures back into the golden dataset so the dev loop regression-tests against them next time.
Canonical instance: Zepto production loop¶
- Effective sample 18–20% (~14,400 traces/day) out of 100,000+ tickets/day.
- Captures 45–60% of edge cases; detects issues in 4–6 minutes.
- Versus uniform sampling: 86% lower review cost per issue identified and 9× better edge-case detection (Source: sources/2026-09-09-databricks-evaluation-first-ai-agents-how-zepto-scales-customer-support).
The article's own summary of the lesson: "Sampling strategy matters more than sampling rate. Naive uniform sampling is cheap but blind. Stratified sampling focused on high-risk flows gives an order of magnitude better issue detection per dollar."
Related instance: the uniform baseline (Figma Response Sampling)¶
Figma's Response Sampling system inspects a configurable fraction of outbound API responses asynchronously for sensitive-data exposure — the same "sample a slice of live traffic and check it out-of-band" shape, but with uniform-random sampling "tuned to balance coverage against overhead" (Source: sources/2026-04-21-figma-visibility-at-scale-sensitive-data-exposure). Figma is the uniform baseline this pattern improves on: the moment failures or exposures concentrate in identifiable strata (a risky endpoint, a new field, a high-privilege user class), risk-weighting the rate would find them faster per unit of inspection cost.
Why it works¶
- Failures are not uniformly distributed. Concentrating the sample where risk concentrates raises detected-issues-per-dollar by roughly an order of magnitude.
- Budget is decoupled from coverage of what matters. You can hold total eval spend flat while approaching full coverage of the high-risk tail.
- Async, non-blocking. Sampling + scoring happen off the request path, so they add assurance without adding user-facing latency.
Tradeoffs and caveats¶
- Stratum definitions can go stale. Risk shifts (new products, new abuse vectors); a fixed stratification silently under-samples the newest failure mode. The strata need continuous revision, ideally driven by where recent production failures landed.
- Blind spots in low-rate strata. The healthy-majority strata are sampled thinly by design; a rare failure there can go unseen longer than under uniform sampling. Mitigate with a small always-on uniform floor under the stratified layer.
- Estimator bias. Aggregate quality metrics computed over a stratified sample must be re-weighted back to the population, or dashboards will over-represent the high-risk strata and look worse (or better) than reality.
- Requires a labeled/annotated trace substrate to define strata at all — a cost that only pays off once volume is high enough that 100% scoring is infeasible.
When to reach for it¶
- Production traffic volume makes 100% evaluation/scoring uneconomic.
- Failures, abuse, or exposures are concentrated in identifiable segments (customer tier, flow age, sentiment, modality, endpoint).
- You already emit per-interaction traces and have scorers (LLM judges + rules) you can run out-of-band.
Related¶
- patterns/snapshot-replay-agent-evaluation — the dev-loop counterpart; stratified sampling curates which production traces become dev-loop snapshots
- patterns/telemetry-to-lakehouse — the trace substrate stratified sampling draws from
- concepts/llm-as-judge — the scorers applied to sampled traces
- concepts/alert-fatigue — tiered alerting over the scored sample
- systems/mlflow — evaluation/monitoring substrate at Zepto
- systems/figma-response-sampling — the uniform-sampling baseline instance