Welcome back to our module on Practical System Design Interviews. In our last lesson, we designed a high-throughput URL shortener. That system was built around a classic request-response model, optimized for a very high ratio of reads to writes. While that pattern is incredibly common, many modern applications—from collaborative tools to financial tickers and live feeds—require the server to push data to the client in real-time, without the client constantly asking for it.
Today, we'll explore the protocols that make this possible. Our goal is to compare real-time communication protocols, specifically WebSockets and long-polling, and their use cases. Understanding the trade-offs between these technologies is crucial for any architect designing interactive systems and is a frequent topic in system design interviews. This knowledge will serve as the direct foundation for our next lesson, where we'll design a scalable chat application.
The Spectrum of Real-Time Communication
The standard HTTP request-response cycle is inefficient for real-time updates. A client would have to repeatedly ask the server, "Is there anything new yet?" This technique, called short polling, generates a lot of network traffic and server load, most of which results in empty responses. It also introduces latency, as an update is only received on the next poll, not when it actually happens.
To solve these issues, engineers developed more sophisticated techniques. Let's explore the three most common ones: long polling, WebSockets, and Server-Sent Events (SSE).

Long Polling: An HTTP-Based Simulation
Long polling is a clever improvement over short polling that uses standard HTTP. Instead of the server responding immediately, it holds the client's request open until it has new data to send.
Here’s the flow:
- The client sends an HTTP request to the server.
- The server doesn't have new data, so it holds the connection open, not sending a response yet.
- When new data becomes available (e.g., another user posts a message), the server sends the data in the response, completing the request.
- The client receives the data and immediately sends another request, restarting the cycle.
- If the connection is held for too long (e.g., 30 seconds) without new data, the server sends a timeout response, and the client, again, immediately reconnects.
This approach creates the illusion of a server push while still using the familiar request-response model.
When To Choose Long Polling vs Websockets for Real-Time Feeds
The team at getstream.io provides a fantastic, code-level explanation of how long polling works.
Read the section on long polling. Focus on the server-side logic: how it checks for data, holds the connection if none is available, and manages timeouts. This aligns well with your backend experience, showing how async request handling on the server makes this pattern efficient.
While much better than short polling, long polling isn't perfect. There's a small latency gap between receiving a response and making a new request, during which events can be delayed. It also carries the overhead of HTTP headers with every message.
WebSockets: True Bidirectional Communication
For true, low-latency, two-way communication, we need a different protocol altogether. This is where WebSockets come in.
A WebSocket connection starts with a special HTTP request from the client, asking the server to Upgrade the connection. If the server agrees, the connection is elevated from the HTTP protocol to the WebSocket protocol. This creates a persistent, full-duplex communication channel over a single TCP connection.
- Persistent: The connection stays open for the life of the session, avoiding the overhead of re-establishing connections.
- Full-duplex: Both the client and the server can send messages to each other at any time, independently.
- Low Overhead: After the initial handshake, message frames are very small (as little as 2 bytes), making it highly efficient for frequent messages.
Networking Essentials for System Design Interviews w/ Ex Meta Senior Manager
In this video, a former Meta engineering manager explains the role of WebSockets in system design.
Watch the segment on WebSockets from bidirectional communication. Pay close attention to his point about WebSockets being inherently stateful. This is a critical architectural consideration, as it means a server must maintain the state of each open connection, which has significant implications for scalability and fault tolerance.
Server-Sent Events (SSE): The One-Way Alternative
Before we do a direct comparison, it's important to know about one more player: Server-Sent Events (SSE). SSE provides a persistent connection where a server can push data to a client. Unlike WebSockets, however, it is unidirectional (server-to-client only).
SSE is built on top of standard HTTP and has some great features, like automatic reconnection. If the client needs to send data to the server, it would do so using a separate, normal HTTP request. This makes SSE an excellent, simple choice for use cases like news feeds, stock tickers, or notification systems where the client is primarily a passive listener.
The Core Comparison: Choosing the Right Tool
The choice between these protocols is a classic system design trade-off. Let's analyze it across several axes that are critical for building scalable systems.
1. Communication Directionality
This is the most fundamental difference.
- Long Polling & SSE: Unidirectional (server-to-client push).
- WebSockets: Bidirectional (client-to-server and server-to-client).
Use Case: If you are building a chat application where users both send and receive messages, you need bidirectional communication, making WebSockets the natural choice. If you are building a dashboard that displays live sports scores, the client only needs to receive updates, making SSE or long polling a simpler and potentially better fit.
2. Latency and Performance
For applications requiring the lowest possible latency, the protocol mechanics matter.
When To Choose Long Polling vs Websockets for Real-Time Feeds
This article does a brilliant job of visualizing the latency differences.
Read the section on latency profiles. The diagram showing the "reconnection gap" in long polling during a burst of events is key. This is the kind of deep analysis expected in a senior-level interview.
As the article explains, long polling can introduce extra latency for messages that arrive in quick succession. WebSockets, with their persistent connection, deliver each message with consistent, single-transit latency. Furthermore, the minimal framing overhead of WebSockets is significantly more efficient than the full HTTP headers required by each long poll message, a difference that becomes critical at high message volumes.
3. Infrastructure and Scalability
This is where your experience as a backend architect is most relevant. The choice of protocol deeply impacts your system's architecture.
When To Choose Long Polling vs Websockets for Real-Time Feeds
Let's continue with the same article to explore the crucial infrastructure trade-offs.
Please read the entire section on scalability and infrastructure. Focus on these key points for a senior engineer: Statefulness and Resources: Understand that despite misconceptions, both protocols require server resources per connection. However, the stateful nature of WebSockets has deeper implications. Load Balancers: Note the critical difference in how L7 load balancers handle these protocols. WebSockets require explicit Upgrade header support, and for stateful connections, a Layer 4 load balancer might be more appropriate, as discussed in the "Networking Essentials" video. Timeout Alignment: Mismatched timeouts across your infrastructure (client, load balancer, proxy, server) are a common source of bugs. Observability: Standard HTTP logs are blind to WebSocket message traffic. You need custom instrumentation.
4. Reliability and Failure Modes
Finally, how do these systems break?
- Long Polling failures are often visible as standard HTTP errors (
504 Gateway Timeout). Recovery can be simple since the client just needs to issue another HTTP request. However, a server restart can cause a "thundering herd" problem, where all clients try to reconnect simultaneously. - WebSocket connections can drop silently (e.g., due to a NAT timeout). You must implement an application-level heartbeat (ping/pong) mechanism to detect and handle these dead connections. On reconnection, you also need a strategy (like using event cursors, similar to long polling) to replay any messages the client missed while disconnected.
Long Polling, Websockets, Server Sent Events - Who Wins? | Systems Design with Ex-Google SWE
This video offers a concise explanation of the "thundering herd" problem, a classic distributed systems challenge.
Watch the segment from the Thundering Herd problem. The solution proposed—adding random jitter to reconnect attempts—is a standard and effective pattern to mitigate this issue.
An Example: Real-Time Auction Service
Let's ground this with a practical example. Imagine designing a real-time auction service.

Here, WebSockets are an excellent choice.
- Bidders need to receive real-time price updates with minimal delay (server-to-client).
- They also need to submit their own bids instantly (client-to-server).
- The architecture reflects the stateful nature of WebSockets. A dedicated pool of
WebSocket Serversmanages the persistent connections. AConnection Registryservice is needed to track which user is connected to which server and subscribed to which auction—this is a classic pattern for managing state in a scalable WebSocket deployment.
Conclusion
We've delved into the critical differences between long polling and WebSockets, with Server-Sent Events as a key alternative. There is no single "best" protocol; the right choice depends entirely on your application's specific requirements.
Key Takeaways:
- Short Polling: Simple but inefficient. Avoid for scalable systems.
- Long Polling: A clever HTTP-based simulation of server push. Good for simple server-to-client updates and for fallbacks, as it works everywhere HTTP does. Its main drawbacks are latency during bursts and header overhead.
- Server-Sent Events (SSE): The modern choice for simple, unidirectional server-to-client push. It's built on HTTP, is lightweight, and has built-in reconnection logic.
- WebSockets: The most powerful option for true bidirectional, low-latency communication. This power comes at the cost of increased complexity in your infrastructure (stateful servers, L4/L7 load balancing) and application logic (heartbeats, manual reconnection, and state synchronization).
As a general rule:
- Need simple server push (feeds, notifications)? Start with SSE.
- Need true two-way interaction (chat, collaboration, gaming)? You need WebSockets.
- Need broad compatibility or a fallback? Long polling is your reliable, if less performant, option.
Now that we have a solid grasp of these protocols and their trade-offs, we are perfectly positioned for our next lesson. We will apply this knowledge to a classic system design interview problem: designing a scalable chat application, which will lean heavily on the capabilities of WebSockets.