Content Ingestion & Podcast Video Incident Report
Spotify's content ingestion and podcast video incident involved a June 24 publishing delay caused by transcoding infrastructure reaching maximum capacity, a backlog forming, and a software bug underutilizing available resources. The incident was exacerbated by a scheduled batch processing job and recent improvements increasing transcoding costs. To resolve the issue, Spotify stopped the batch job, deployed a fix for the resource scheduling bug, and added additional processing capacity overnight. Following the incident, Spotify increased transcoding capacity by 67%, fixed the resource scheduling bug, and improved monitoring to alert earlier to capacity approaching limits. The company is also investing in broader improvements to podcast publishing pipeline reliability, including better capacity planning, prioritization, and rate limiting mechanisms.