JUL 24, 2026
EngBrief
Search⌘K
LatestTopicsSourcesSaved
Eng&Brief

Engineering insights from the world's best tech companies, curated and summarized.

Weekly brief

Browse

TopicsSourcesFavorites

More

SearchRSS Feed
© 2026 EngBriefUpdated every 4 hours
← Sources
netflix.com icon
Streaming

Netflix TechBlog

44 articles on EngBrief

The Netflix TechBlog is one of the most influential engineering blogs in the industry. Netflix engineers share deep dives into streaming at scale, content delivery, microservices architecture, data engineering, machine learning for recommendations, and the resilience patterns that keep the platform running for 200M+ subscribers worldwide.

StreamingMicroservicesData EngineeringMachine LearningResilience
Visit blog →

Latest Articles

Netflix6d ago

In-House LLM Serving at Netflix

Here's a 2-3 concise sentence summary of the engineering blog post "In-House LLM Serving at Netflix": Netflix implemented an in-house LLM serving system, running the full stack from model deployment to inference, to handle the complexities of large-scale AI workloads and ensure custom models with non-traditional logic could be easily supported. The system uses the vLLM engine and Triton Inference Server, with a Java control plane for deployment, versioning, and health checking, allowing for seamless model upgrades and experimentation-to-production paths. By using an OpenAI-compatible API interface and reusable client libraries, the platform simplifies the migration of LLM models from hosted to self-hosted environments.

StreamingScale
12 min
Netflix10d ago

Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned

Here's a 3-sentence summary of the engineering blog post "Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned" from Netflix TechBlog: To build a reliable and real-time service dependency map at Netflix scale, engineers adopted a "streaming-first" architecture that continuously ingests flow records and provides near-real-time topology updates with backpressure-enabled reactive pipelines to handle massive scale without data loss. A multi-layer architecture with physically separate topology layers and independent optimization enabled the system to process millions of flow records per second and provide sub-second query responses while maintaining near-real-time freshness. Through lessons learned, Netflix engineers found that backpressure and physical storage isolation are crucial components for building reliable real-time systems at scale, and that solving complex engineering challenges requires embracing system complexity and trade-offs.

StreamingScale
25 min
Netflix25d ago

GenPage: Towards End-to-End Generative Homepage Construction at Netflix

Authors: Lequn Wang, Jiangwei Pan, and Linas BaltrunasFigure 1. Autoregressive homepage generation. GenPage builds a Netflix homepage one row or entity at a...

StreamingScale
25 min
NetflixJun 23, 2026

Toward More Controllable AI Video Editing: An Early Research Exploration at Netflix

By Zhuoning Yuan, Ta-Ying Cheng, Benjamin Klein, Bahareh AzarnoushIntroductionAt Netflix, we build technology to help storytellers bring their creative visions...

StreamingScale
11 min
NetflixJun 22, 2026

How Netflix Simplified Batch Compute with Kueue

By Alvin Bao, Alex Petrov, Jennifer Lai, Aidan Sherr, and Samartha ChandrashekarAs a part of the journey to transition Netflix’s compute infrastructure to be...

StreamingScale
7 min
NetflixJun 19, 2026

The Data Canary: How Netflix Validates Catalog Metadata

By Celina AmadosAt Netflix, our catalog metadata is crucial to our member experience, and a single corrupted data state can impact millions of viewers...

StreamingScale
8 min
NetflixJun 19, 2026

Data Projects: Managing Data Assets at Netflix Scale

By Amer Hesson, Marcelo Mayworm, James Mulcahy, and Brittany TruongThe Problem: Managing Assets at Netflix ScaleNetflix’s Data Platform is vast. We have...

StreamingScale
10 min
NetflixJun 19, 2026

Predicting Risk in Content Launches: How Data-Driven Insights can Transform Launch Planning

by Emily GillEach year, we bring the Analytics Engineering community together for an Analytics Summit — a multi-day internal conference to share analytical...

StreamingScale
8 min
NetflixJun 19, 2026

The Evolution of Cassandra Data Movement at Netflix

By Guil Pires, Jennifer Prince, Jose Camacho, Ken Kurzweil, Phanindra ChunduruBackgroundIn a previous post, we introduced Data Bridge, a unified management...

StreamingScale
14 min
NetflixJun 19, 2026

Thinking Fast & Slow for a Personalized Notification System

by Matthew Wood, Ishan Gupta, Kevin Mercurio, Devon Bryant, and Claire DormanIn his seminal book “Thinking, Fast and Slow,” Daniel Kahneman describes two...

StreamingScale
7 min
NetflixJun 19, 2026

A Human-Augmenting Agentic Workflow for Causal Inference

By Winston Chou, Adrien Alexandre, Lars Olds, Yi Zhang, Garrett Hagemann, and Nathan KallusIntroductionImagine asking a data agent to analyze the causal...

StreamingScale
13 min
NetflixJun 19, 2026

VMAF v1: Good Is Not Good Enough

By Christos G. Bampis, Zhi Li, Kyle Swanson, Nil Fons Miret and Pavan MadhusudanaraoWill this encode look good to Netflix members? Does switching to a new...

StreamingScale
12 min
NetflixJun 19, 2026

VMAF v1: Good Is Not Good Enough

By Christos G. Bampis, Zhi Li, Kyle Swanson, Nil Fons Miret and Pavan MadhusudanaraoWill this encode look good to Netflix members? Does switching to a new...

StreamingScale
12 min
NetflixJun 8, 2026

A Human-Augmenting Agentic Workflow for Causal Inference

By Winston Chou, Adrien Alexandre, Lars Olds, Yi Zhang, Garrett Hagemann, and Nathan KallusIntroductionImagine asking a data agent to analyze the causal...

StreamingScale
13 min
NetflixJun 5, 2026

Thinking Fast & Slow for a Personalized Notification System

by Matthew Wood, Ishan Gupta, Kevin Mercurio, Devon Bryant, and Claire DormanIn his seminal book “Thinking, Fast and Slow,” Daniel Kahneman describes two...

StreamingScale
7 min
NetflixJun 3, 2026

Dynamically Splitting Wide Partitions in Cassandra for Time Series Workloads

By Rajiv Shringi, Kaidan Fullerton, Oleksii Tkachuk and Kartik SathyanarayananIntroductionNetflix’s TimeSeries Abstraction is a scalable system for ingesting...

StreamingScale
12 min
NetflixJun 3, 2026

Dynamic Repartitioning for Time Series Workloads

By Rajiv Shringi, Kaidan Fullerton, Oleksii Tkachuk and Kartik SathyanarayananIntroductionNetflix’s TimeSeries Abstraction is a scalable system for ingesting...

StreamingScale
12 min
NetflixMay 29, 2026

High-Throughput Graph Abstraction at Netflix: Part I

By Oleksii Tkachuk, Kartik Sathyanarayanan, Rajiv ShringiIntroductionNetflix has a diverse range of graph use cases, each serving specific business needs with...

StreamingScale
14 min
NetflixMay 29, 2026

From Silos to Service Topology: Why Netflix Built a Real-Time Service Map

By Parth Jain, Rakesh Sukumar, Yingwu Zhao, Renzo Sanchez & Nathan FisherHow we built a living map of our distributed infrastructure to help engineers...

StreamingScale
15 min
NetflixMay 29, 2026

From Silos to Service Topology: Why Netflix Built a Real-Time Service Map

By Parth Jain, Rakesh Sukumar, Yingwu Zhao, Renzo Sanchez & Nathan FisherHow we built a living map of our distributed infrastructure to help engineers...

StreamingScale
15 min