Skip to main content
Create your own

DNS Load Balancing: Round Robin & Weighted Records

Hello! Welcome to the fourth lesson in our module on Network Infrastructure and DNS.

In our last lesson, we analyzed the critical role of DNS TTL, establishing that low TTLs are a prerequisite for any agile DNS-based strategy. We saw how TTL creates a direct trade-off between performance and the ability to make rapid changes.

Today, we build directly on that foundation to address the learning outcome: Implement DNS-based load balancing using round-robin and weighted records, analyzing its limitations for high-load systems.

We will explore how DNS can be used for simple load distribution. However, given your background in building resilient, high-performance systems, our primary focus will be a critical analysis of why this technique, despite its simplicity, is often inadequate for modern, high-load applications.

1. DNS Round Robin: The Simplest Form of Load Distribution

DNS Round Robin is a basic load balancing technique where multiple IP addresses are associated with a single hostname. When a DNS query is made for that name, the DNS server returns the full list of IPs but rotates their order for each subsequent request.

The idea is that different clients will receive the IPs in a different order and will typically connect to the first one in the list, thereby distributing the load across the servers.

What is Round Robin Load Balancing? Definition & FAQs

To start, let's get a clear, concise definition of Round Robin DNS from this VMware article.

Please read the sections 'What is Round Robin Load Balancing?', 'How Does Round Robin Load Balancing Work?', and 'What is the Difference Between Round Robin DNS vs. Load Balancing?'. Focus on how the DNS server itself performs the load balancing by rotating IP addresses.

To see this mechanism in action, including some of its real-world complexities, the following video provides a practical demonstration.

DNS Round Robin and Netmask Ordering

This video from ITFreeTraining demonstrates how to configure and test DNS Round Robin. It also introduces some interesting edge cases that we will dissect later.

Please watch these segments: Introduction (00:00 - 02:04): This explains the core concept of cycling through DNS records. Configuration and Testing (04:49 - 09:40): Pay attention to how multiple A records are created for the same name and how nslookup is used to observe the rotation. Note the important point that you must use A records, not CNAMEs, for this. We will return to the other parts of this video later.

Implementation is straightforward: for a hostname like api.example.com, you would create multiple A/AAAA records in your DNS zone:

Name Type Value
api.example.com A 192.0.2.1
api.example.com A 192.0.2.2
api.example.com A 192.0.2.3

2. Weighted Round Robin: Distributing Load Unevenly

A significant drawback of standard Round Robin is that it assumes all servers are identical. In reality, you often have a heterogeneous fleet of servers with varying capacities. Weighted Round Robin addresses this by allowing you to assign a "weight" to each record, directing a proportional amount of traffic to each server.

This is not a standard DNS protocol feature but rather a service offered by modern DNS providers like AWS Route 53, Google Cloud DNS, and others.

What is Round Robin Load Balancing? Definition & FAQs

The VMware article we looked at earlier also provides a good introduction to the concept of weighted distribution.

Read the section 'What is the Difference Between Weighted Load Balancing vs Round Robin Load Balancing?'. It clearly illustrates how weights translate to request distribution.

For a practical look at how this is configured in a real-world system, the AWS Route 53 documentation is an excellent example.

Weighted routing

This documentation for AWS Route 53 shows how weighted routing is implemented as a managed DNS feature.

Read the introductory section on 'Weighted routing'. Note how traffic is proportioned based on the ratio of a record's weight to the total weight of all records in the group. This is a common pattern for A/B testing or canary releases.

For example, to send 80% of traffic to a new, powerful server cluster and 20% to an older one, you might configure weights of 80 and 20, or simply 4 and 1. The absolute numbers don't matter, only their ratio.

3. Critical Analysis: Why DNS Load Balancing Fails High-Load Systems

While simple to set up, DNS-based load balancing has severe limitations that make it unsuitable for most high-availability, high-load services. Your experience with low-latency and payment systems has likely exposed you to the need for much more sophisticated traffic management.

Limitation 1: The Caching Problem and Slow Failure Detection

This is the most significant flaw. As we discussed in the last lesson, DNS responses are cached aggressively throughout the internet.

  • Scenario: A server at 192.0.2.2 fails.
  • DNS Server: The DNS server itself has no idea the server is down. It continues to include 192.0.2.2 in its round-robin responses.
  • Client/Resolver: A client's recursive resolver has a cached response that includes the dead IP. Even if you set a low TTL of 60 seconds, clients will continue to receive and attempt to connect to the failed IP for up to 60 seconds.
  • Impact: For that entire duration, a fraction of your users will experience connection timeouts and errors. A 60-second recovery time is unacceptable for any critical service.

Dedicated load balancers, by contrast, use active health checks to detect a failed backend in seconds and immediately remove it from the pool.

Limitation 2: No Awareness of Server Health or Load

The DNS server is completely blind to the state of the application servers. It doesn't know if a server is:

  • Offline or unreachable.
  • Overloaded with high CPU or memory usage.
  • Healthy but returning application-level errors (e.g., HTTP 500s).

It will mechanically send traffic to a dead or struggling server just as readily as a healthy one. Some advanced DNS services (like AWS Route 53) attempt to solve this by integrating health checks, which can automatically remove records for unhealthy targets.

Weighted routing

Let's revisit the AWS documentation to see how they address this limitation.

Read the section 'Health checks and weighted routing'. This shows how a proprietary DNS feature can add health awareness, turning DNS into a slightly more intelligent system. However, this is an add-on, not a core feature of DNS itself.

Even with health checks, the recovery is still bound by the DNS TTL, making it slower than a dedicated load balancer.

Limitation 3: Uneven Load Distribution due to Caching

The "round-robin" distribution is often a theoretical ideal. In practice, large groups of users (e.g., an entire corporation or university) are often behind a single set of recursive DNS resolvers.

When that resolver queries for api.example.com, it gets a list of IPs (e.g., [192.0.2.2, 192.0.2.3, 192.0.2.1]). It caches this specific response and serves it to all users behind it for the duration of the TTL. The result is that thousands of users from that organization will all be directed to a single server (192.0.2.2), completely defeating the load balancing.

4. Valid Use Cases for DNS-Based Load Balancing

Despite these serious drawbacks, DNS load balancing is not useless. It is a valid tool for coarse-grained, global-level traffic management.

  1. Global Server Load Balancing (GSLB): This is its most powerful and common use case. Instead of balancing traffic between individual servers in a single datacenter, you balance traffic between entire datacenters. For example, www.example.com could resolve to:

    • The public IP of the load balancer in your US-East datacenter (for users in North America).
    • The public IP of the load balancer in your EU-West datacenter (for users in Europe).
      This is typically combined with GeoDNS (which we'll cover next) to direct users to the closest datacenter.
  2. Simple Datacenter Failover: Using weighted records, you can implement a simple DR strategy.

    • Primary DC (US-East): Weight 100
    • DR DC (US-West): Weight 0
      In a disaster, you flip the weights (e.g., to 0 and 100) to redirect all traffic to the DR site. The recovery is slow (bound by TTL), but it's a simple and effective mechanism for catastrophic failures.

Conclusion

This lesson explored DNS-based load balancing, a technique that is simple in principle but fraught with practical limitations for high-load systems.

Key Takeaways:

  • DNS Round Robin distributes requests by rotating the order of IP addresses in DNS responses.
  • Weighted Round Robin allows for proportional traffic distribution, useful for heterogeneous servers or canary deployments.
  • The primary limitations are its inability to react quickly to failures due to DNS caching, its lack of awareness of server health/load, and its tendency to create uneven load distribution.
  • Its most valid modern use case is for coarse-grained Global Server Load Balancing (GSLB), directing traffic between geographically distinct datacenters, not between individual servers within one.

Preview of the next lesson:
We've established that for fine-grained, highly available load balancing, DNS is not the right tool. This naturally leads us to dedicated load balancers. However, before we configure proxies like nginx, we need to understand the layer they operate on. In the next module, we will begin by analyzing the TCP connection lifecycle in high-load systems, which is fundamental to understanding how modern load balancers achieve high performance and efficiency.

Can't find a good explanation? Sign up and we'll make it for you

Sign up