Skip to content

SPOTIFY

Read original ↗

Spotify — Content Ingestion & Podcast Video Incident Report

Summary

A public incident report on the June 24, 2026 podcast video publishing delay at Spotify, plus the broader reliability program now underway across the podcast publishing pipeline. When a creator publishes an episode, the audio/video goes through processing steps — transcoding (converting media into app-ready formats) and content analysis — before becoming available. On June 24 the video transcoding infrastructure hit maximum capacity, building a queue backlog that delayed video-podcast publication for several hours (normally minutes). Four factors converged: (1) transcoding ran with insufficient headroom to absorb large spikes of bulk content delivery; (2) a scheduled batch re-processing job was consuming capacity alongside real-time new-episode processing; (3) a resource-scheduling bug following a migration to more powerful hardware left compute underutilized by ~10%; and (4) creators, getting no confirmation their upload was received, re-uploaded, adding further load. Remediation: stop the batch job, deploy the scheduling fix, add processing capacity overnight; backlogs cleared by the next morning. Follow-ups target capacity planning for burst + recovery, stronger prioritization of real-time content over background work, pipeline-wide rate limiting and backpressure, and earlier creator notification.

Key takeaways

  1. The trigger was capacity exhaustion under a bulk-delivery spike, not a hard outage. Video transcoding "reached maximum capacity," and a queue built up: medium-priority (new episodes) and low-priority (updates to older episodes) queues both grew, then drained once capacity was freed and added. The system stayed up but latency to publish blew out from minutes to hours. (Source: this article.) See concepts/elasticity.

  2. Insufficient headroom for spikes was the root capacity failure. Transcoding could scale to handle typical submission periods and background low-priority work, but "there wasn't sufficient headroom to scale and support large spikes caused by the bulk delivery of new content." The permanent fix increased transcoding capacity by ~67% to restore the missing headroom. See concepts/elasticity and overnight-capacity-add-to-drain-backlog.

  3. A batch job contended with the real-time path. A routine job re-processing existing episodes (to stay compatible with playback-system changes) "appeared fine earlier in the day" but "became problematic when combined with increased content submissions" — background work stealing capacity from creator-facing publication. Stopping it (at 16:35 UTC) was the first mitigation. See concepts/oltp-vs-olap and priority-queue-real-time-over-batch.

  4. Priorities existed but weren't strong enough. The pipeline already distinguishes medium-priority (new episodes) from low-priority (updates to older episodes) queues — yet a batch job and a spike still starved real-time publishing. A named follow-up is "improving prioritization across our publishing systems so that real-time content from creators is always processed ahead of background operations." See priority-based-queue-scheduling.

  5. A resource-scheduling bug silently cut throughput ~10%. After migrating to more powerful hardware, a bug in resource scheduling caused systems to underuse available processing capacity, reducing throughput by about 10%. This is a classic post-migration efficiency regression: the hardware was present but the scheduler didn't place work on it. See resource-scheduling-underutilization.

  6. Missing upload acknowledgment amplified load. Because the system did not confirm uploads were received and queued, creators who saw nothing appear re-uploaded episodes — adding load during the exact window the pipeline was already saturated. Spotify explicitly owns this: "the system should have confirmed their upload was received and queued, and it did not." See upload-acknowledgment-before-processing.

  7. Detection lagged the failure by ~4 hours — a time-to-detection gap. Early alerts fired at 13:30 UTC but were "not immediately recognized as a broader capacity issue." Engineers stopped the batch job at 16:35, but the full scope wasn't recognized until queue-backlog thresholds were breached at 17:34 and incident response formally began. The monitoring follow-up ("alert earlier when capacity is approaching limits") exists to close that gap. See time-to-detection.

  8. The remediation was three coordinated actions. Stop the batch job (free capacity), deploy the scheduling-bug fix (recover the lost ~10%), and add a new processing cluster overnight (raise ceiling). All queues cleared by 01:02 UTC Jun 25; full confirmation at 07:30.

Operational numbers

  • ~67% — permanent increase in transcoding capacity after the incident, to restore headroom for spikes + batch operations.
  • ~10% — throughput lost to the resource-scheduling bug that underused newly migrated, more-powerful hardware.
  • Queue tiers: medium priority = new episodes; low priority = updates to older episodes (both backed up and drained during the incident).
  • ~4 hours between first alerts (13:30 UTC) and formal incident response (17:34 UTC).

Timeline (UTC, Jun 24 unless noted)

Time Event
13:30 Early alerts fire in internal monitoring; not recognized as a capacity issue.
15:00 Video-podcast delivery spike pushes transcoding near maximum capacity.
16:35 Batch processing job stopped to free capacity.
17:31 First creator report of podcast-video publishing impact.
17:34 Automated alerts confirm queue backlog exceeding thresholds; incident response begins.
19:00 Creator reports escalated to incident team.
20:49 Software fix deployed to improve resource utilization.
00:14 (Jun 25) Additional processing cluster brought online.
01:02 All queues cleared.
07:30 Full confirmation: all publishing pipelines operating normally.

Caveats

  • This is a creator-reliability / publishing-latency incident report, not a deep architecture paper. It names the pipeline stages (transcoding, content analysis) and queue priorities but does not detail the transcoding cluster internals, the scheduler, or the batch system.
  • The forward-looking reliability program (capacity planning, prioritization, rate limiting/backpressure, creator notification) is described as intent, not yet as shipped design.

Source

Last updated · 766 distilled / 2,225 read