Scaling Reliable Experimentation in a Two-Sided AdTech Marketplace: ZMS Budget Split¶
Summary¶
Zalando Marketing Services (ZMS) runs a retail-media / sponsored-products marketplace: advertisers compete for ad placements, customers see the results, and the whole system is bounded by a finite shared resource — advertisers' campaign budgets. This post explains why the standard user-level A/B test breaks in this setting and how ZMS made experimentation reliable and concurrent between 2023 and 2025 using two capabilities: Budget Split (allocate a separate budget pool per experiment variant) and Orthogonal Concurrency (partition budgets into 2ⁿ orthogonal buckets so multiple experiments can run at once without interfering). The central diagnosis is a SUTVA violation the post calls Cannibalization Bias: when Treatment and Control draw from the same campaign budget, a more efficient variant wins more auctions and drains the shared wallet faster, starving the other variant's funding and artificially depressing its metrics — so Control is no longer a valid counterfactual.
Key takeaways¶
-
Retail-media experimentation is a two-sided-marketplace problem, not a consumer-app problem. The standard A/B assumption "what User A does doesn't affect User B" fails because advertisers compete for placements through a finite shared budget (Source: sources/2026-08-31-zalando-scaling-reliable-experimentation-in-a-two-sided-adtech-marketplace).
-
Cannibalization Bias is a concrete SUTVA violation through a shared wallet. Consider a novel bidding algorithm (Variant A) vs. baseline (Variant B) drawing from one campaign budget: A's superior efficiency wins more auctions and depletes the budget faster, which restricts B's access to funding and artificially depresses B's metrics. This "dilutes the true average treatment effect and skews the experiment's results."
-
Budget Split creates "parallel universes." Instead of Treatment and Control fighting over one pool, ZMS divides the campaign budget upstream proportional to the traffic split, so each variant runs against its own isolated budget — effectively two identical, independent marketplaces. If Treatment beats Control, it's because of the feature, not because it cannibalized the other variant's budget.
-
The mechanism has three stages under the hood: (a) traffic randomization — assign users to Treatment/Control; (b) budget partitioning — split every participating ad campaign into virtual "sub-campaigns"; (c) isolated auctioning — Treatment users only trigger bids from the Treatment sub-campaign budget, Control users only from the Control sub-campaign budget.
-
Orthogonal Concurrency generalises Budget Split to concurrent experiments. If two experiments (each A/B) run simultaneously, a campaign's budget is split into four orthogonal buckets: 1A2A, 1A2B, 1B2A, 1B2B. In general, N concurrent A/B experiments require 2ⁿ buckets. Campaigns starting at different times get their budgets dynamically partitioned across combinations of variants from multiple overlapping experiments.
-
Absolute isolation is impossible; ZMS monitors for residual SUTVA violations. Four named residual channels remain even after Budget Split:
- Algorithmic campaign steering — daily budget shifts and max-bid adjustments optimise on combined Treatment+Control performance, so a treatment that moves run-rate or ROAS indirectly shifts budget/bids for Control.
- Pre-experiment ML effects — pCTR/pCVR models train daily on the last 7 / 28 days; pre-experiment event distributions (reflecting the control status quo) leak into the treatment group's training distribution even with separate training pipelines.
- Manual advertising-operations interventions — ops staff adjust budgets/bids across markets on aggregate performance, a human-in-the-loop interference channel.
-
Timing / right-censoring — campaigns extending beyond the experiment window create right-censoring selection effects that skew revenue interpretation.
-
When isolation can't be guaranteed, pivot to causal inference. "When A/B testing is not possible, or if we cannot guarantee isolation, we pivot to causal inference methodologies (like non-experimental counterfactual analysis)" — the same design stance as Zalando's Octopus platform, which ships quasi-experimental tooling rather than forcing A/B on every use case.
-
Impact — reliability and concurrency both unlocked. Before Budget Split, experiments were "flaky and had to be done sequentially" — a reliability problem and a concurrency problem (tests fought for the same traffic and budget). After: ZMS experimentation scaled from 8 experiments (2023, no Budget Split) to more than 60 experiments (2025) with the isolation properties above, increasing volume while enabling more frequent and more reliable decisions. Budget Split is framed as "a prerequisite for trust" for PMs, analysts, and applied scientists — it lets them "fail faster when hypotheses don't pan out."
Operational numbers¶
- 8 experiments in 2023 (without Budget Split) → 60+ experiments in 2025 (with Budget Split + Orthogonal Concurrency).
- 2ⁿ budget buckets for N concurrent A/B experiments (worked example: 2 experiments → 4 buckets: 1A2A / 1A2B / 1B2A / 1B2B).
- ML model training windows: 7 days (pCTR) and 28 days (pCVR), retrained daily.
Systems / concepts / patterns extracted¶
- Systems: Zalando Marketing Services (ZMS) retail-media / sponsored-products marketplace and its experimentation framework; sibling to Octopus.
- Concepts: budget-cannibalization-bias, budget-split, orthogonal-concurrency, sutva (violation), user-split-experiment (fails here), market-mediated-long-term-effects (adtech variant), quasi-experimental-methods (the fallback), causal-inference, sample-ratio-mismatch.
- Patterns: per-variant-resource-isolation (the generalised pattern Budget Split instantiates), centralized-experimentation-platform, patterns/staged-rollout.
Caveats¶
- The post is a capabilities/architecture narrative, not a paper: it discloses the mechanism (traffic randomization → budget partitioning → isolated auctioning) and the 2ⁿ-bucket scheme but not the internal service architecture of the auction/budgeting system.
- The only hard numbers are the 8→60+ experiment count and the ML training windows; no dollar figures, statistical-power numbers, or SRM rates are disclosed for ZMS specifically.
- Budget Split reduces but does not eliminate interference — the post is explicit that four residual SUTVA-violation channels remain and are actively monitored, and that causal inference is the fallback when isolation can't be guaranteed.
Source¶
- Original: https://engineering.zalando.com/posts/2026/09/scaling-reliable-experimentation-in-two-sided-adtech-marketplace.html
- Raw markdown:
raw/zalando/2026-08-31-scaling-reliable-experimentation-in-a-two-sided-adtech-marke-8fec62ad.md
Related¶
- companies/zalando
- user-split-experiment — the A/B shape that fails here.
- market-mediated-long-term-effects — the Lyft-anchored marketplace-interference sibling.
- quasi-experimental-methods — the fallback when isolation can't be guaranteed.
- systems/octopus-zalando-experimentation-platform — Zalando's other experimentation platform.