Skip to content

SYSTEM Cited by 2 sources

Spotify Publishing Pipeline

Definition

The Spotify publishing pipeline (a.k.a. content-ingestion pipeline) is the set of processing stages a podcast episode goes through, after a creator hits publish, before it becomes available to listeners. Named stages (Source: sources/2026-07-20-spotify-content-ingestion-podcast-video-incident-report):

  • Transcoding — converting uploaded audio/video media into the formats Spotify's apps need. For video podcasts this runs on dedicated video transcoding infrastructure that autoscales with submission volume.
  • Content analysis — additional processing on the media.

Work flows through priority-tiered queues:

  • Medium priority — newly published episodes (the real-time, creator-facing path).
  • Low priority — updates/re-processing of older episodes (background work, e.g. re-encoding existing content to stay compatible with playback-system changes).

Under normal operation, video podcast episodes publish within minutes.

How it behaves under load

The pipeline is elastic but not infinitely so. During the June 24 2026 incident the transcoding tier hit maximum capacity during a bulk-delivery spike, and the medium- and low-priority queues both backed up — turning a minutes-long publish into a multi-hour delay. The pipeline stayed available; what degraded was publish latency (a latency, not availability failure mode).

Contributing weaknesses surfaced by that incident:

  • No headroom for spikes. Transcoding could scale for typical submissions plus background work, but not for large bursts of bulk new content. Fixed by a ~67% capacity increase (concepts/elasticity).
  • Batch/real-time contention. A scheduled re-processing (batch) job consumed capacity alongside new-episode processing (concepts/oltp-vs-olap); priorities existed but did not strongly guarantee real-time-over-background (priority-based-queue-scheduling).
  • Scheduler underutilization. After a hardware migration, a resource- scheduling bug left ~10% of compute unused (resource-scheduling-underutilization).
  • No upload acknowledgment. The pipeline didn't confirm receipt/queueing to creators, who re-uploaded and amplified load (upload-acknowledgment-before-processing).

Reliability program (stated follow-ups)

  • Capacity planning for steady-state plus burst capacity plus incident recovery.
  • Stronger prioritization so real-time creator content is always processed ahead of background operations (priority-queue-real-time-over-batch).
  • Rate limiting and backpressure extended throughout the pipeline to handle unexpected load gracefully.
  • Earlier creator notification when publishing isn't working.

Follow-through remediations (2026-09-16)

A later Spotify reflection frames the two structural weaknesses behind the June incident as pre-existing to the content spike, and reports the shipped fixes (Source: sources/2026-09-16-spotify-ai-changed-how-spotify-builds-quality-at-higher-velocity):

  • Silent failures could be masked. A media file that couldn't be processed sometimes failed silently and paged no one, so publishing impact went unnoticed for hours. Fix: end-to-end monitoring so failures are known before creators notice — the observability gap, not just the capacity gap.
  • No burst capacity for the growing video catalog. Valid video episodes could queue unalerted when transcoding capacity was exhausted. Fix: increase capacity and, critically, fix the scheduler and move batch jobs to lower priority so background work yields to new uploads.
  • Service tiering + workload prioritization reworked so critical services and new uploads take precedence when capacity is constrained — the pipeline instance of criticality-based load shedding (recorded as prose; folded into concepts/graceful-degradation).
  • Bad-actor suppression as load control. Episodes from bad actors are suppressed and greatly de-prioritized, cutting overall load so they don't compete with higher-priority episodes — prioritization doubling as a load-shedding lever.

Spotify's note: AI helped deliver these fixes faster, but they required "distinct judgement, an end-to-end mindset and skills from our engineers" — the failures were "garden variety … and ours," not AI-authored.

Seen in

Last updated · 766 distilled / 2,225 read