Welcome to the next lesson in our module on microservices. In our last session, you successfully built a gRPC service and client in Go, mastering the foundation of high-performance internal communication. We concluded by asking a critical question: how do external clients, like web and mobile apps, communicate with our internal, gRPC-based microservices?
Today, we answer that question by introducing the API Gateway. This architectural pattern serves as the single, managed entry point to your entire system. Your goal for this lesson is to understand how to configure an API Gateway to handle three of its most critical responsibilities: request routing, authentication using JWTs, and rate limiting. Mastering this component is essential for building secure, resilient, and scalable distributed systems.
The Role of the API Gateway
Before diving into configuration, let's solidify our understanding of what an API Gateway is and why it's a cornerstone of modern microservice architectures. At its core, an API Gateway is a specialized reverse proxy that sits between your clients and your backend services. It abstracts the complexity of the internal service landscape and provides a unified, secure interface to the outside world.
The short video "What is API Gateway?" from ByteByteGo provides a quick, high-level overview of its purpose and functions.
Watch this entire short video to get a concise summary of an API Gateway's role and the typical flow of a client request.
Watch the full video from start to finish. Pay attention to the list of functions it handles, such as authentication, rate limiting, and protocol translation (e.g., HTTP to gRPC), which directly relates to the services you built in the previous lesson.
The "API Gateways 101" diagram provides a more detailed visual breakdown of these concepts.

Notice a few key points from this diagram:
- Decoupling: Clients talk to one endpoint (the gateway), not dozens of individual services. This simplifies client-side code and allows you to refactor or re-platform internal services without impacting clients.
- Centralized Features: The gateway handles "cross-cutting concerns" like authentication, rate limiting, and routing. This keeps your microservice code clean and focused on business logic. Without a gateway, each service would need to implement this logic independently, leading to duplication and inconsistency.
Now, let's configure these features in practice. We'll focus on Kong, a popular open-source API Gateway, because of its powerful plugin architecture and straightforward declarative configuration.
Configuring an API Gateway with Kong
The video "Kong Gateway Tutorial" provides an excellent, hands-on demonstration of setting up Kong. It uses Docker Compose and a declarative YAML file (kong.yml) to define the gateway's behavior, which is a common and effective approach in production environments.
First, let's focus on the two most fundamental tasks: defining backend services and creating routes to them.
- Service: A service is an abstraction that represents an upstream backend API or microservice (e.g., your
user-serviceororder-service). - Route: A route defines rules for how requests are matched and forwarded to a service. Most commonly, this is based on the request's path (e.g.,
/usersmaps to theuser-service).
Kong Gateway Tutorial | API Gateway For Beginners
This video walks you through setting up Kong with Docker and defining services and routes in a declarative kong.yaml file.
Watch the segment from setting up the demo. Focus on how the kong.yaml file is structured. Observe how two services are defined—one for an external API (GitHub) and one for a local service. Then, see how routes are created with specific paths (/gists and /hello) to direct traffic to the correct service. The final curl commands demonstrate that the routing works as configured.
As you can see, basic request routing is simple yet powerful. You can now imagine mapping different paths to the various gRPC services you might build, with the gateway handling the protocol translation from HTTP to gRPC.

Securing Services with Authentication (JWT)
A public-facing API must be secured. A gateway is the perfect place to enforce authentication, ensuring that no unauthenticated traffic reaches your internal services. We will focus on JSON Web Tokens (JWT), a standard for stateless authentication.
The authentication flow is as follows:
- A client authenticates with an identity provider (e.g., a login service) and receives a JWT.
- The client includes this JWT in the
Authorization: Bearer <token>header for all subsequent requests to the API Gateway. - The gateway intercepts the request, validates the JWT's signature and expiration, and checks its claims.
- If the token is valid, the gateway forwards the request to the appropriate backend service, often adding headers with user information (like
X-User-ID) extracted from the token. If invalid, the gateway rejects the request with a401 Unauthorizederror.
The "Securing Backend with API Gateways" section of the "API Gateways 101" diagram illustrates this flow perfectly. Now, let's look at how this is implemented.
The article "How to Build API Gateway Architecture" provides excellent examples of how to configure this. Since you are proficient in JavaScript, the Node.js middleware example will be very clear.
How to Build API Gateway Architecture
This article section explains JWT authentication and provides a practical code example.
In the article, find section 5. Authentication and Authorization. Read the subsection on JWT Authentication. The Node.js code shows exactly what the gateway needs to do: extract the token, verify it using a secret, and handle errors like expiration. This is the logic that a gateway's JWT plugin executes internally.
In Kong, you don't need to write this code yourself. You simply enable and configure the JWT plugin. The article "How to Configure Kong Ingress Rate Limiting and Authentication Plugins" shows the declarative YAML configuration for this.
How to Configure Kong Ingress Rate Limiting and Authentication Plugins
This resource demonstrates configuring Kong plugins in a Kubernetes environment, but the YAML for the plugin itself is universally applicable.
Find the section Configuring Authentication Plugins and review the YAML configuration for JWT Authentication. Notice how you can declaratively specify which claims to verify (exp, nbf) and the maximum token expiration.
By centralizing JWT validation at the gateway, your backend services no longer need to worry about it. They can trust that any request they receive has already been authenticated.
Protecting Services with Rate Limiting
The final core function we'll cover is rate limiting. This is crucial for protecting your backend services from being overwhelmed by traffic, whether from malicious actors (DDoS attacks), buggy clients, or simply a surge in legitimate usage.
The gateway tracks the number of requests from a given client (identified by IP address, API key, or user ID) and blocks them if they exceed a configured threshold (e.g., 100 requests per minute).
The "Kong Gateway Tutorial" video provides a simple, clear demonstration of adding a rate-limiting plugin.
Kong Gateway Tutorial | API Gateway For Beginners
This part of the video shows how to add and configure a rate-limiting plugin for a specific service.
Watch the final segment from adding a plugin. Observe how a plugins block is added to the kong.yaml file. The rate-limiting plugin is enabled for the hello-service and configured for 5 requests per minute. The subsequent test with curl clearly shows the gateway allowing the first 5 requests and then blocking the 6th with a 429 Too Many Requests error.
This simple configuration is incredibly powerful. For a truly scalable system, you would configure the rate-limiting plugin to use a distributed cache like Redis to share request counts across multiple gateway instances. The article "How to Build API Gateway Architecture" discusses this advanced pattern.
How to Build API Gateway Architecture
This section provides a more advanced implementation of rate limiting, suitable for distributed systems.
In section 6. Rate Limiting, review the Node.js Rate Limiter Implementation. This example uses a sliding window algorithm with Redis. This is a common interview topic and a pattern used in high-scale systems to ensure rate limits are enforced correctly across a cluster of stateless gateway nodes. Note the policy: redis setting in the Kong configuration example, which achieves the same result declaratively.
Conclusion
In this lesson, you've learned how an API Gateway acts as the front door to a microservices architecture. By centralizing key functionalities, it simplifies your backend services, enhances security, and improves resilience.
Key Takeaways:
- Single Entry Point: The API Gateway provides a single, stable endpoint for all external clients, decoupling them from the internal service architecture.
- Centralized Cross-Cutting Concerns: Logic for routing, authentication, and rate limiting is handled at the gateway, keeping microservices lean and focused on their core domain.
- Declarative Configuration: Modern gateways like Kong allow you to define complex routing, security, and traffic control rules in simple, version-controlled YAML files.
- Core Functions: You now understand how to configure the three most important gateway functions:
- Routing: Mapping URL paths to specific backend services.
- Authentication: Offloading JWT validation to protect internal services.
- Rate Limiting: Protecting services from traffic spikes and abuse.
We've touched upon several rate-limiting strategies and algorithms (like Token Bucket and Sliding Window). In our next lesson, we will dive deeper into this topic, comparing the most common rate limiting algorithms and analyzing their trade-offs, which is a frequent subject in system design interviews.