Skip to content

PATTERN Cited by 2 sources

Stratified evaluation sampling

Stratified evaluation sampling is the production-loop pattern of scoring only a fraction of live traffic, with per-stratum sampling rates weighted by risk and value instead of a single uniform rate — so that a fixed evaluation budget is concentrated on the interactions most likely to be wrong, costly, or dangerous.

Intent

Scoring 100% of production traffic with LLM judges and rule checks is prohibitively expensive at scale. The naive alternative — uniform random sampling at some low rate (e.g. 10%) — is cheap but blind: because failures and abuse concentrate in a small subset of traffic, uniform sampling spends most of its budget re-confirming the easy, healthy majority and misses most edge cases.

Stratification fixes the allocation: partition traffic into strata by risk/value and sample each stratum at a rate proportional to how much you need to see it. The total sample size stays within budget, but its composition shifts toward the interactions that matter.

Mechanism

  1. Instrument every interaction as a trace (the substrate — see patterns/telemetry-to-lakehouse / patterns/snapshot-replay-agent-evaluation).
  2. Define strata by risk and value. At Zepto the sampling rate is raised for: high-value customers; new or recently-changed flows; negative sentiment / high escalation risk; and image-based or fraud-prone interactions (Source: sources/2026-09-09-databricks-evaluation-first-ai-agents-how-zepto-scales-customer-support).
  3. Assign per-stratum sampling rates so the effective overall rate stays within the cost envelope while high-risk strata approach full coverage.
  4. Score the sampled traces with the same scorers that gate deployment (LLM judges + deterministic rules), write results to a governed store, and drive dashboards + tiered alerts (alert rules).
  5. Feed detected failures back into the golden dataset so the dev loop regression-tests against them next time.

Canonical instance: Zepto production loop

The article's own summary of the lesson: "Sampling strategy matters more than sampling rate. Naive uniform sampling is cheap but blind. Stratified sampling focused on high-risk flows gives an order of magnitude better issue detection per dollar."

Figma's Response Sampling system inspects a configurable fraction of outbound API responses asynchronously for sensitive-data exposure — the same "sample a slice of live traffic and check it out-of-band" shape, but with uniform-random sampling "tuned to balance coverage against overhead" (Source: sources/2026-04-21-figma-visibility-at-scale-sensitive-data-exposure). Figma is the uniform baseline this pattern improves on: the moment failures or exposures concentrate in identifiable strata (a risky endpoint, a new field, a high-privilege user class), risk-weighting the rate would find them faster per unit of inspection cost.

Why it works

  • Failures are not uniformly distributed. Concentrating the sample where risk concentrates raises detected-issues-per-dollar by roughly an order of magnitude.
  • Budget is decoupled from coverage of what matters. You can hold total eval spend flat while approaching full coverage of the high-risk tail.
  • Async, non-blocking. Sampling + scoring happen off the request path, so they add assurance without adding user-facing latency.

Tradeoffs and caveats

  • Stratum definitions can go stale. Risk shifts (new products, new abuse vectors); a fixed stratification silently under-samples the newest failure mode. The strata need continuous revision, ideally driven by where recent production failures landed.
  • Blind spots in low-rate strata. The healthy-majority strata are sampled thinly by design; a rare failure there can go unseen longer than under uniform sampling. Mitigate with a small always-on uniform floor under the stratified layer.
  • Estimator bias. Aggregate quality metrics computed over a stratified sample must be re-weighted back to the population, or dashboards will over-represent the high-risk strata and look worse (or better) than reality.
  • Requires a labeled/annotated trace substrate to define strata at all — a cost that only pays off once volume is high enough that 100% scoring is infeasible.

When to reach for it

  • Production traffic volume makes 100% evaluation/scoring uneconomic.
  • Failures, abuse, or exposures are concentrated in identifiable segments (customer tier, flow age, sentiment, modality, endpoint).
  • You already emit per-interaction traces and have scorers (LLM judges + rules) you can run out-of-band.
Last updated · 766 distilled / 2,225 read