Skip to main content
Create your own

Optimizing TCP Connections in High-Load Systems

Hello! Welcome to your first lesson in the "High-Load Distributed Systems" course.

Given your extensive background in building high-performance systems, we'll be moving quickly from foundational principles to the practical, in-depth details you're looking to master.

Introduction

This lesson kicks off our module on Load Balancing and Reverse Proxies. We'll begin at the most fundamental layer of network communication that underpins all distributed systems: the TCP connection.

Our learning outcome for this session is to analyze the TCP connection lifecycle in high-load systems, identifying connection reuse and pooling optimizations.

We will cover:

  • A brief review of the TCP connection lifecycle, focusing on states critical to performance.
  • The practical limits on TCP connections and common bottlenecks like ephemeral port exhaustion.
  • Key optimization strategies: connection reuse (HTTP Keep-Alive) and connection pooling.

Understanding how to manage and optimize TCP connections is the bedrock for building scalable and resilient load balancers, database clients, and microservice communication, which we will explore in subsequent lessons.

Let's get started.


1. The TCP Connection Lifecycle and Its Limits

While the three-way handshake (SYN, SYN-ACK, ACK) and four-way teardown are familiar concepts, in high-load systems, the nuances of connection states and resource limits become critical.

A server can theoretically handle a vast number of connections, limited mainly by memory and CPU. However, a single client connecting to a single server is constrained by the number of available source ports, known as ephemeral ports.

To understand these limits, please watch the following video. It clearly explains the distinction between server-side capacity and the client-side port limitation, which is a crucial concept, especially when a reverse proxy acts as a client to your backend services.

Is there a Limit to Number of Connections a Backend can handle?

This video from Hussein Nasser's channel explains the theoretical and practical limits of TCP connections.

Please watch the first three sections (timestamps 00:41 to 10:11). Focus on: The 16-bit port header and the resulting ~65,000 connection limit per client IP. How a server can handle millions of connections from different clients. Why a reverse proxy, especially at Layer 4, can become a bottleneck by exhausting this connection limit when communicating with backend servers.

As the video highlights, the 4-tuple that uniquely identifies a TCP connection is (source_ip, source_port, destination_ip, destination_port). When a proxy or any client application makes many connections to the same destination service, it can run out of available source_ports, a condition known as ephemeral port exhaustion.

This problem is often exacerbated by connections lingering in the TIME_WAIT state after being closed. This state is designed to ensure any delayed packets from the old connection are properly handled, but it holds the port, making it unavailable for a period (typically 30-120 seconds).

For a deep, practical dive into how these limits manifest and how to tune the underlying OS, the following article is an excellent resource. It details a real-world attempt to scale a server to 1 million connections.

TCP & Journey of achieving 1 Million connections — Part 1

This article, 'TCP & Journey of achieving 1 Million connections — Part 1', provides a hands-on look at OS-level TCP tuning.

Please read the sections 'Important configs' and 'Road to 100K TCP connections'. Focus on these key points: net.ipv4.ip_local_port_range: Understand how this kernel parameter defines the pool of ephemeral ports. net.ipv4.tcp_tw_reuse and net.ipv4.tcp_fin_timeout: See how these settings control the TIME_WAIT state to reclaim ports faster. The tcpkali Debugging Scenario: Pay close attention to the analysis of TIME_WAIT connections. It's a perfect example of how an application-level setting (IdleTimeout) directly caused TCP-level bottlenecks by prematurely closing connections, leading to port exhaustion.

Now that we've established the costs and limits associated with creating and closing connections, let's explore the primary optimization techniques.

2. Optimization 1: Connection Reuse (HTTP Keep-Alive)

The most basic optimization is to stop creating a new connection for every request. The overhead of the TCP handshake, especially in high-latency networks, is significant.

HTTP Keep-Alive (or persistent connections) is a mechanism that allows the same TCP connection to be used for multiple, sequential HTTP requests and responses. This avoids the repeated cost of connection setup and teardown.

This diagram illustrates HTTP Keep-Alive. A single TCP connection is established between the client and server, which is then reused to transfer multiple files (CSS, JS, images) instead of creating a new connection for each one.

On the server side, a crucial parameter is the idle timeout. This setting determines how long the server will keep an idle connection open before closing it. As we saw in the tcpkali example from the previous resource, setting this value requires a careful trade-off:

  • Too short: You lose the benefit of reuse, as connections are closed too quickly.
  • Too long: You risk tying up server resources (memory, file descriptors) with idle connections, which can be a vector for denial-of-service attacks.

3. Optimization 2: Connection Pooling

Connection pooling takes the concept of reuse a step further. Instead of just reusing a single connection, a connection pool manages a "pool" or cache of pre-established, open TCP connections that can be borrowed by application threads, used, and then returned for others to use.

This is a standard pattern for managing database connections, but the principle applies to any scenario involving frequent communication with another service, such as a microservice calling another via HTTP or gRPC.

Boost Your Go App's Network Performance with a TCP ...

This article, 'Boost Your Go App's Network Performance with a TCP Connection Pool', provides an excellent conceptual overview and design considerations for connection pools.

Please read the introduction and the first two main sections ('Core Concepts' and 'Designing a TCP Connection Pool'). Focus on: The core workflow: Initialize, Borrow, Use, Return. The key design goals: Speed, Reliability, Scalability. The critical configuration parameters: Max Connections, Min Idle Connections, and Timeouts. The importance of health checks to ensure connections in the pool are still valid.

The performance impact of connection pooling is not just theoretical. It provides a dramatic reduction in latency and an increase in throughput by eliminating the handshake delay for most requests.

To see a practical demonstration with hard numbers, the following video compares the performance of a simple API with and without database connection pooling.

Connection Pooling in PostgresSQL with NodeJS (Performance Numbers)

This video provides a clear, practical demonstration of the performance gains from connection pooling.

Watch the introduction (00:00 - 01:08) to understand the two approaches being compared. Then, skip to the performance demonstration (09:34 - 11:21). Notice the significant difference in response time (~50% improvement) between the 'old' method (new connection per request) and the 'pool' method.

In your work with low-latency trading and high-load payment systems, you've undoubtedly leveraged connection pooling for database access. The principles we've discussed here—managing pool size, idle timeouts, and health checks—are universal and apply directly when configuring HTTP clients for microservice communication or setting up proxies like nginx and Envoy, which we will cover soon.


Conclusion

In this lesson, we analyzed the fundamental layer of high-load system communication.

Key Takeaways:

  • The TCP connection lifecycle has performance-critical states, with TIME_WAIT being a primary cause of ephemeral port exhaustion.
  • While a server's connection capacity is limited by resources, clients (including proxies) are constrained by the ~65,000 ephemeral ports available per IP address.
  • Connection Reuse (HTTP Keep-Alive) is a baseline optimization that amortizes the cost of the TCP handshake over multiple requests.
  • Connection Pooling is a more advanced application-level pattern that manages a set of ready-to-use connections, dramatically reducing latency and controlling resource consumption. Understanding and tuning pool parameters (max connections, idle timeout) is critical for scalability.

Preview of the Next Lesson:

A reused or pooled connection is only useful if it's healthy. What happens if a firewall silently drops an idle connection, or a server crashes? Your application might try to use a "dead" connection, leading to long timeouts and errors. In our next lesson, we will address this by learning how to configure TCP keepalive parameters to detect dead connections in load balancers.

Can't find a good explanation? Sign up and we'll make it for you

Sign up