AI Changed How Spotify Builds: What We Learned (and Fixed) About Quality at Higher Velocity¶
Summary¶
Spotify Engineering's reflection on running quality and reliability while AI-assisted development has roughly doubled the volume of merged change year-over-year. The headline finding is counterintuitive: across incident retrospectives Spotify did not find AI-authored code to be a material direct contributor to incidents. The real strain is second-order — the volume of change grew faster than the verification controls (review, testing, rollout, observability, rollback, failover) could adapt. The post walks four production areas each tested at once — content processing, fleet updates, compute shortages / regional failover, and mobile app release quality — and the concrete remediations for each, then presents Spotify's own data on whether velocity is trading off against quality (it argues not). The recurring theme: AI increased the capacity to produce change; the next constraint became the ability to verify it.
Operating context (scale numbers, all from this post): 777M monthly active users across 2,000+ device types, ~100M concurrent clients, 11–12M backend requests/sec, ~3,000 production services, and >500K new songs/videos/podcasts/ audiobooks ingested per day.
Key takeaways¶
-
No distinct AI-authored failure signature. Spotify added two questions to its monthly major-incident retrospective during the AI ramp-up: (a) did AI-authored code directly contribute? (b) did the increased volume of change pressure review, testing, rollout, or observability? Across incidents reviewed, (a) was not material; (b) was observed — "the volume of change increased faster than some of our verification controls could adapt." The response is to strengthen the entire delivery system, not to blame the model. (Source: sources/2026-09-16-spotify-ai-changed-how-spotify-builds-quality-at-higher-velocity)
-
Content processing: silent failures + no burst headroom. The content-ingestion pipeline had two pre-existing weaknesses: processing failures could be masked (an unprocessable media file failed silently, paging no one, so publishing impact went unnoticed for hours), and the pipeline lacked capacity for spikes in the growing video catalog (valid episodes queued unalerted when transcoding capacity was exhausted). This is the same class of failure as the June 24 incident (see sources/2026-07-20-spotify-content-ingestion-podcast-video-incident-report).
-
Content-processing fixes. Added end-to-end monitoring so failures are known before creators notice; fixed the scheduler; moved batch jobs to lower priority; increased capacity; and reworked service tiering + workload prioritization so critical services and new uploads take precedence when capacity is constrained. Episodes from bad actors are suppressed and de-prioritized, cutting overall load so they don't compete with higher-priority episodes. (Source: sources/2026-09-16-spotify-ai-changed-how-spotify-builds-quality-at-higher-velocity)
-
Fleet updates: automation creates new failure modes. Spotify's custom Fleet Management framework makes large-scale fleet changes daily, mostly auto-merged after safety checks, now extended to agentic-driven changes (a recent Java migration across backend services finished in three days). But an automated dependency upgrade passed the checks and still failed in production, impacting end users. Response: strengthen safeguards, expand rollback capacity, and schedule automated changes during the owning team's working hours (so a human is present when a fleet change lands). (Source: sources/2026-09-16-spotify-ai-changed-how-spotify-builds-quality-at-higher-velocity)
-
Compute shortages made regional failover user-visible. Industry-wide AI demand spiked CPU/GPU demand without matching supply, shrinking spare capacity. Spotify historically had compute available on demand; now it's less predictable. When a region fails, Spotify shifts traffic to another region — a regional failover engineered to minimize customer impact. Earlier this year, the capacity shortage exacerbated previously-trivial failover issues and end users noticed. This is framed as an indirect AI impact on quality of service, not the commonly-cited direct one.
-
Edge / tiering rework for the capacity-constrained world. Spotify is reviewing its network edge and tiering approach: on failover it must now accept there may be no capacity for lower service tiers — i.e. explicitly shed lower-priority work to protect critical paths (graceful degradation under constrained capacity). It doubled reserved edge capacity after a May incident. Today it can shift part of internal service-mesh traffic manually; extending that control to edge traffic and testing regional spillover are still under way. Goal: move traffic gradually while ensuring the receiving region can absorb the load and stay stable. (Source: sources/2026-09-16-spotify-ai-changed-how-spotify-builds-quality-at-higher-velocity)
-
Mobile app quality is a cyclical guardrail problem. Over a decade Spotify has seen a ebb-and-flow: ship features fast → quality signals degrade → shift capacity/incentives to quality → broaden guardrail metrics → recover → new issues emerge outside existing metrics → repeat. AI raises the frequency of this cycle, so gaps surface faster. The insight: individual releases can look healthy while smaller regressions accumulate over time, affect particular phones, or sit outside watched signals. Fix: broaden quality signals and add longer-term trend analysis to release decisions to catch deterioration earlier.
-
The data argues against a simple quality-for-velocity trade-off. Merged changes more than doubled YoY in August (~8,100 → 17,000). Spotify classifies every merged PR (features / code-quality-and-optimization / maintenance / documentation): quality & optimization rose from 27% → 31% of the mix — meaning >2× the absolute quality work YoY. Maintenance/config fell 31% → 25%. Google Cloud's 2025 DORA research found AI adoption correlated with higher throughput and product performance but lower delivery stability; Spotify chose to answer the question from its own data rather than assume.
-
Rework rate stayed flat despite the industry churn spike. Spotify rebuilt its rework-rate metric to separate genuine rework from new work and legacy refactoring. Code churn = code removed relative to added; rework rate weighs the age of code being changed (a better proxy for whether recent work holds up). The FAROS 2026 report found a sharp industry-wide rise in code churn, but Spotify sees no corresponding rise in rework rate — a signal it is not accumulating AI-induced quality debt.
-
Two warning signals, deliberately not "fixed" yet. Code complexity and PR size are both creeping up. Pre-AI these were unambiguous quality concerns; now a larger PR may just mean a human + agent reasoned together and safely delivered a bigger unit of work, and complexity thresholds calibrated for one human's head may no longer apply. Spotify has no conviction on either hypothesis, so it deliberately refuses to relax the thresholds to feel better and keeps watching whether they're truly leading indicators. (Source: sources/2026-09-16-spotify-ai-changed-how-spotify-builds-quality-at-higher-velocity)
Systems / concepts / patterns extracted¶
- Systems: Fleet Management / Fleetshift (agentic fleet changes, dependency-upgrade prod failure, rollback/working-hours scheduling remediation), content-ingestion pipeline (silent-failure + burst-capacity fixes), Honk (the agent behind fleet mutations).
- Concepts: regional failover (new page — canonical, cited across the corpus), graceful degradation (shed lower tiers on constrained failover), elasticity / headroom (burst + failover headroom), blast radius (schedule automated fleet changes during owning-team hours so a human is present), observability (end-to-end failure monitoring + longer-term trend signals), backpressure (workload controls / priority under saturation), OLTP vs OLAP (real-time uploads must beat background batch).
- Patterns: fast rollback (expand rollback capacity as a safeguard for higher change volume), service tiering / criticality-based load shedding (recorded as prose + tag — see below), staged rollout (move traffic gradually on failover/spillover).
Taxonomy notes (why so few new pages)¶
Per the AGENTS.md canonicalization gate:
concepts/regional-failover— MINTED. It's the textbook term, is cited across 13+ existing source pages (AWS multi-region DR, Meta zGateway, Atlassian StreamHub, PlanetScale consensus, Redpanda Shadow Link, Airbnb monitoring, Databricks Lakebase, …) with no dedicated page, so it clears the ≥3-citation promotion bar comfortably.- service tiering / criticality-based load shedding — NOT minted. Recorded as a
tagand prose here + folded into concepts/graceful-degradation. It's close to canonical but the wiki already routes "shed low-priority work to protect critical paths" through graceful degradation; left for Lint promotion if it earns its own citations.patterns/service-tieringis referenced in the body as a wikilink so Lint can promote it later if it recurs. - verification-gap / quality-at-velocity, capacity-planning, rework-rate-vs-code-churn, guardrail-metrics — NOT minted. These are this-article framings / metric definitions and single-source, recorded as tags + prose only.
Operational numbers¶
| Metric | Value | Note |
|---|---|---|
| Monthly active users | 777M | across 2,000+ device types |
| Concurrent clients | ~100M | steady-state |
| Backend requests/sec | 11–12M | |
| Production services | ~3,000 | |
| Daily content ingested | >500K | songs/videos/podcasts/audiobooks |
| Merged changes (Aug YoY) | ~8,100 → 17,000 | more than doubled |
| Quality/optimization share of PRs | 27% → 31% | >2× absolute quality work |
| Maintenance/config share of PRs | 31% → 25% | |
| Java fleet migration duration | 3 days | agentic Fleet Management |
| Reserved edge capacity | 2× | doubled after a May incident |
| Fleet scheduler bug (June incident) | ~10% throughput loss | referenced from July incident report |
Caveats¶
- This is a reflection / methodology post, not an incident postmortem — it references the June 24 content-ingestion incident (separately reported) rather than re-deriving it.
- The "no AI-authored failure signature" finding is Spotify-specific and self-reported, explicitly contrasted against DORA 2025 (which found lower delivery stability with AI) and FAROS 2026 (which found rising code churn). Spotify's counter-evidence (flat rework rate, rising quality-work share) is presented as their data, not an industry conclusion.
- Edge-traffic control and regional-spillover testing are described as still in progress, not shipped.
- Complexity and PR-size trends are flagged as unresolved — Spotify explicitly declines to conclude they are (or aren't) quality problems.
Source¶
- Original: https://engineering.atspotify.com/2026/9/ai-changed-how-spotify-builds-what-we-learned-and-fixed-about-quality-at-higher-velocity/
- Raw markdown:
raw/spotify/2026-09-16-ai-changed-how-spotify-builds-what-we-learned-and-fixed-abou-97e02f0b.md
Related¶
- companies/spotify
- systems/spotify-fleetshift
- systems/spotify-publishing-pipeline
- systems/spotify-honk
- concepts/regional-failover
- concepts/graceful-degradation
- concepts/elasticity
- concepts/blast-radius
- concepts/observability
- patterns/fast-rollback
- sources/2026-07-20-spotify-content-ingestion-podcast-video-incident-report