PATTERN Cited by 2 sources
Experimental-tier model promotion¶
Definition¶
Experimental-tier model promotion is the lifecycle for adopting a third-party frontier model of unknown quality into a large user population: give everyone access immediately on an explicitly-labelled experimental tier, cap the blast radius with a dedicated per-user "experimental" budget, gather multi-signal evidence for a few days, then promote the model into standard circulation (or make it the default) or drop it.
It is the model-adoption analogue of a staged rollout, but the artefact being rolled out is not your own code or config — it is an external vendor's model whose marketed "frontier" status may be false. The pattern exists because two things are simultaneously true at scale:
- Speed matters. New models ship almost weekly and the efficiency frontier advances fast; waiting to evaluate before granting access forfeits the largest available cost lever.
- Naive adoption is dangerous both ways. A model marketed as frontier can regress quality (Databricks: Opus 5.0 ranked lower than 4.8 while costing more) and explode cost (a no-mitigation control-group rollout of "GPT Astra" raised average developer spend 60% overnight).
The resolution is to decouple access (instant, for everyone) from endorsement (earned, after evidence) — with a budget bucket bounding the cost of being wrong during the evaluation window. (Source: sources/2026-09-28-databricks-how-databricks-rolls-out-frontier-models-to-12000-employees)
The three stages¶
- Immediate experimental access. The moment a model releases, expose it to all users behind an "Experimental" tag so they can select it but know it "may or may not be best in class or stick around forever." This requires a central control plane that can flip a model on server-side and a client-side distribution mechanism to push the new model config to every user's tool (Databricks pushes via the UG CLI over Mobile Device Management — see patterns/progressive-configuration-rollout).
- Budget-constrained exposure. The experimental model draws from a dedicated per-user experimental budget — a fraction of the monthly budget set aside for untested models — so widely exposing an off-frontier model has a bounded cost. This is a distinct budget bucket from the daily-runaway and monthly-maximum limits, and from a separate quality-frontier budget that rations the most premium models (which are known-good but 2–3× the cost of the next tier, so meant for specialized tasks, not daily driving). See systems/unity-ai-gateway-budgets.
- Promote or drop on multi-signal evidence. After a short window (Databricks: ~3 days), fuse three signals to decide: private benchmarks (offline task suites + online side-by-side runs on real work like PR creation — concepts/llm-as-judge), user-reported quality (power users compare over Slack/surveys), and cost tracking (per-session cost from OpenTelemetry traces — patterns/telemetry-to-lakehouse). Public benchmarks are treated as insufficient (concepts/benchmark-methodology-bias). Outcomes are not binary "keep/kill": a model can be promoted to default, promoted to GA but not default, added only to a smart router's toolkit for its cost profile (concepts/model-first-routing), or dropped from the catalog.
Preconditions¶
- An AI gateway that can add / remove / tag models centrally and meter per-model, per-user spend across providers (patterns/unified-billing-across-providers). Without cheap model-switching, Day-1 access isn't affordable.
- Per-user budget buckets that can be scoped to a model class (experimental / quality-frontier / standard), not just a global cap.
- A representative internal evaluation — offline + online — because the whole point is that vendor/public "frontier" claims don't predict your workload. Capturing the frontier requires knowing your incumbents.
- Per-session cost comparison that survives distribution shift. Early
adopters try harder problems with new models, so a raw
$/sessionaverage is misleading; sessions must be stratified (e.g. single- vs multi-turn × file-edits-or-not) and re-weighted before comparison — a cost-side sibling of patterns/stratified-evaluation-sampling.
Why not just staged rollout?¶
patterns/staged-rollout and patterns/progressive-configuration-rollout gate exposure of your artefact by cohort (canary → region → fleet) on health signals. Experimental-tier model promotion differs on three axes:
- Exposure is fleet-wide from day one, not cohort-staged — the blast radius is bounded by a budget, not by user count.
- The artefact is external and adversarially-marketed — the decision is procurement (is this vendor's claim true for us?), not release safety.
- The gating signal is quality-vs-cost on real work, not error-rate/latency health. A model that is perfectly healthy (no errors) can still be dropped for being off the efficiency frontier.
The config-distribution half of the pattern is progressive configuration rollout (pushing the new model config out through the gateway + laptop CLI); the evaluation-and-promotion half is what this pattern adds.
Seen in¶
- sources/2026-09-28-databricks-how-databricks-rolls-out-frontier-models-to-12000-employees — canonical instance. Databricks gives 12,000+ employees Day-1 experimental access to new models via Unity Gateway + UG CLI, caps them with a per-user experimental budget (one of four budgets, alongside a new quality-frontier budget), and promotes/drops in ~3 days on benchmarks + user reports + OTel cost tracking. Sept-21 launch week: Opus 5 / GPT-6 Sol / Luna → GA by Day 3; Opus 5.5 slated as Claude Code default; GPT-6 Sol added to the smart router (not made default). Stratified $/session comparison showed −29% (Opus) and −48% (GPT).
- sources/2026-08-07-databricks-managing-ai-coding-costs-at-scale — the precursor framing: a central model menu on the gateway + internal automated evals to decide which new models beat incumbents (GLM rolled out on a positive eval; Opus/Stripe negative evals declined). Establishes the "evaluate-then-roll-models" half without the explicit experimental-budget mechanism.
Related¶
- concepts/efficiency-frontier — the promote/drop decision criterion.
- concepts/model-first-routing — a promotion outcome: add to the router's toolkit.
- concepts/benchmark-methodology-bias — why private, not public, benchmarks decide.
- patterns/progressive-configuration-rollout — the config-distribution half (push new model config to every client).
- patterns/staged-rollout — sibling for cohort-staged code exposure (contrast: budget-bounded here).
- patterns/ai-gateway-provider-abstraction — the control plane that makes cheap model-switching possible.
- patterns/unified-billing-across-providers — the metering substrate the budgets ride on.
- patterns/stratified-evaluation-sampling — the re-stratification method reused for cost comparison.
- systems/unity-ai-gateway / systems/unity-ai-gateway-budgets — the productised instance.
- companies/databricks.