Welcome back! In our last session, we designed a scalable chat application, where you saw how a publish/subscribe system could "fan out" messages to manage group chats. This concept of delivering a single piece of information to many recipients is a fundamental pattern in distributed systems.
Today, we'll explore this pattern in much greater depth by tackling another quintessential system design interview problem: designing a news feed system. Our goal is to design a system that employs a fan-out architecture for message distribution. This is a crucial topic for your goal of mastering scalable systems, as it forces a direct confrontation with one of the most classic trade-offs in the field: optimizing for writes versus optimizing for reads. We will dissect the strategies used by giants like Twitter and Facebook to serve billions of feeds daily.
Step 1: Defining the Scope
As you prepare for senior-level interviews, you know that the first step is always to clarify the problem. A news feed can be incredibly complex, with algorithmic ranking, ads, and various content types. To make this tractable for a one-hour session, we'll focus on the core functionality.
The ByteByteGo guide provides a concise, interview-style dialogue for establishing the scope of a news feed system. This is an excellent model for how you should approach this phase in a real interview.
Please read the first section, "Step 1 - Understand the problem and establish design scope". This will help us define the core functional and non-functional requirements for our design.
Based on this, let's establish our requirements:
-
Functional Requirements:
- A user can publish a post.
- A user can follow other users.
- A user can see a news feed containing posts from the people they follow.
- The feed should be sorted in reverse chronological order (newest first).
-
Non-Functional Requirements:
- Low Latency on Read: Viewing the news feed should be very fast, ideally under a few hundred milliseconds.
- High Availability: The system must be resilient to failures.
- Eventual Consistency: It's acceptable if a new post takes a short time (e.g., up to a minute) to appear in all followers' feeds.
With these requirements, the central challenge becomes clear: how do we efficiently deliver a post from one user to potentially millions of followers while keeping feed retrieval fast?
Step 2: The Fundamental Trade-Off: Push vs. Pull
At the heart of any news feed design lies a choice between two architectural models for distributing content. This is often framed as "fan-out on write" vs. "fan-out on read".
- Fan-out on Write (Push Model): When a user creates a post, the system immediately "pushes" that post (or a reference to it) into the feed of every single follower. The work is done at write time. This means each follower's feed is pre-computed and ready to be served.
- Fan-out on Read (Pull Model): When a user wants to see their feed, the system "pulls" the latest posts from all the people they follow, merges them, sorts them, and then presents the result. The work is done at read time.
Your experience with small-scale systems likely involved the pull model, as it's simpler to implement when the number of followed accounts is small. However, at scale, the choice has massive implications for performance and cost.
Re-evaluating Fan-Out-on-Write vs. Fan-Out-on-Read Under Celebrity Traffic Spikes (2025)
This technical analysis from Codemia provides an excellent deep dive into the trade-offs between these two models. It's a more advanced take than typical interview prep material and perfect for understanding the underlying principles.
First, read the section "The Basics" for a clear definition of each model. Then, carefully study Table 1, which provides a summary of the key trade-offs between the push, pull, and hybrid approaches. This table crystallizes the entire problem domain.
As you can see from the reading, the choice is not simple:
- The Push model optimizes for fast reads but creates enormous write amplification.
- The Pull model optimizes for simple writes but makes reads slow and computationally expensive.
Step 3: The Push Model in Detail: Fan-out on Write
Let's design a system based on the fan-out-on-write (push) model, as our primary non-functional requirement is low latency on reads. This architecture prioritizes the user experience of scrolling the feed.
The high-level flow when a user creates a post looks like this:

Let's break down this asynchronous process. Because you've worked with Go and have an interest in resilient systems, think of this as a set of interacting goroutines or microservices communicating via a message broker like RabbitMQ or Kafka.
Design FB News Feed System Design Interview w/ ex: Meta Senior Manager
This video segment from Hello Interview clearly explains how to implement the fan-out-on-write model using an asynchronous worker pool. It also introduces the critical "celebrity problem" and the hybrid solution we will discuss next.
Watch from the "what can we do instead?" section. Focus on these key architectural components: The use of a precomputed feed table/cache to store each user's feed. The introduction of an asynchronous worker pool and a message queue to handle the fan-out process without blocking the post creation request. The idea of breaking large fan-out jobs into smaller sub-jobs to distribute the load.
This asynchronous fan-out is the core of the architecture. The key components are:
- Post Service: Receives the new post, stores it in a
Post DB, and publishes apost_createdevent to a message queue. - Message Queue: Decouples the write path from the fan-out logic, ensuring that even if fan-out is slow, the user who posted gets a fast response.
- Fan-out Workers: A pool of services that consume events from the queue. For each
post_createdevent, a worker:
a. Retrieves the author's follower list from aUser/Social Graph DB.
b. For each follower, it inserts thepost_idinto their respective feed list in aNews Feed Cache. A system like Redis is ideal here, where each user's feed can be a list or sorted set keyed byuser_id.
When a user requests their feed, the Feed Service simply reads the list of post_ids from the News Feed Cache for that user, a very fast operation.
Step 4: The Celebrity Problem and the Hybrid Solution
The pure fan-out-on-write model has a glaring weakness, often called the "celebrity problem" or "hot publisher problem." What happens when a user with 50 million followers posts? Your fan-out workers suddenly have to perform 50 million writes to the News Feed Cache. This can create a massive "thundering herd" of activity that can overwhelm your cache, database, and network.
This is where the industry-standard hybrid model comes in.
System Design Interview Walkthrough: Design Twitter
This video on designing Twitter directly addresses the celebrity problem and explains the hybrid approach that Twitter itself adopted.
Please watch the section covering the timeline service, specifically from the "one catch" for celebrity users. It concisely explains the hybrid fan-out strategy.
The hybrid approach elegantly solves the problem:
- For most users (e.g., those with < 10,000 followers), continue using the fan-out-on-write (push) model.
- For "celebrity" users (those with a very high follower count), do not fan out their posts on write. Instead, treat them as a special case.
When a user requests their feed, the process becomes:
- Fetch the pre-computed feed of
post_idsfrom theirNews Feed Cache(this contains posts from all the non-celebrities they follow). - Identify the list of celebrities the user follows.
- Fetch the most recent posts directly from those celebrities (a fan-out-on-read/pull operation).
- Merge the two lists of posts in the Feed Service, sort chronologically, and return the result.
This hybrid model contains the write storm from celebrities while still providing the low-latency read benefits of the push model for the vast majority of user interactions.
Step 5: The Read Path and Feed Hydration
So far, we've only discussed lists of post_ids. A client application needs much more: the post content, author's name and profile picture, like/comment counts, etc. The process of fetching this full data is called hydration.

The read path for our hybrid system looks like this:
This section from ByteByteGo details the retrieval and hydration process.
Read the section "Newsfeed retrieval deep dive". Pay close attention to step 5, which explains how the service fetches complete objects from user and post caches to construct the "fully hydrated" feed.
As the resource describes, the News Feed Service acts as an aggregator on the read path. It:
- Retrieves the list of
post_ids(after merging celebrity and non-celebrity posts). - For each
post_id, it fetches the full post data from aPost Cache(falling back to thePost DB). - For each post's
author_id, it fetches the user data from aUser Cache(falling back to theUser DB). - It assembles the final JSON response and sends it to the client.
Extensive caching at each layer (News Feed Cache, Post Cache, User Cache) is absolutely critical to meeting our low-latency requirement.
Conclusion
In this lesson, you designed a scalable news feed system by navigating the critical trade-offs of a fan-out architecture. This problem beautifully illustrates the challenges you'll face when moving from small-scale applications to systems that serve millions.
Key Takeaways:
- Central Trade-off: News feed design is a classic battle between fan-out-on-write (push) for fast reads and fan-out-on-read (pull) for simple writes.
- The Celebrity Problem: A pure fan-out-on-write model creates massive write spikes for popular users, a major bottleneck at scale.
- The Hybrid Solution: The industry-standard solution is a hybrid model. Use the push model for the majority of users and switch to the pull model for celebrities. This balances read performance with write-path stability.
- Asynchronous Processing: Using message queues and background workers is essential for a resilient and scalable write path, decoupling the user-facing action from the heavy lifting of the fan-out.
- Hydration and Caching: The read path involves not just fetching post IDs but also "hydrating" them with full content, a process that relies heavily on a multi-layered caching strategy to remain fast.
You have now designed two of the most common systems asked about in interviews: a chat application and a news feed. In our final lesson, we will focus on how to synthesize and articulate these complex designs using a structured framework, preparing you to confidently present your architectural vision in a high-stakes interview.