Skip to main content
Create your own
Lesson illustration

SLIs and SLOs: Defining and Tracking Service Levels

Welcome to the next lesson in our journey through system observability. In the previous session, we configured Nginx to act as a reverse proxy and load balancer. This gave us a scalable and resilient way to distribute traffic to our backend services. However, simply distributing traffic isn't enough. How do we know if our services are actually healthy and meeting user expectations? How do we make objective, data-driven decisions about whether to release a new feature or fix a stability issue?

This lesson introduces the foundational concepts of Site Reliability Engineering (SRE) that answer these questions: Service Level Indicators (SLIs) and Service Level Objectives (SLOs). You will learn to define a common, quantitative language for service reliability. Mastering this framework is a critical step toward your goal of operating high-scale systems, as it provides the mechanism to balance innovation with the reliability that keeps users happy.

The Language of Reliability: SLI, SLO, and SLA

Before we dive deep, it's crucial to understand the terminology. These three acronyms are related but distinct.

  • Service Level Indicator (SLI): A quantitative measure of some aspect of your service's performance. Think of it as a specific metric that reflects the user's experience. An SLI is typically expressed as a ratio of "good events" to "total valid events."
  • Service Level Objective (SLO): A target value or range for an SLI, measured over a specific time window. This is the reliability goal you set for your service.
  • Service Level Agreement (SLA): A formal contract with your users that specifies consequences (like refunds or service credits) if your SLOs are not met.

In engineering practice, our focus is on SLIs and SLOs. They are the internal tools we use to build and maintain reliable systems. SLAs are business and legal constructs that follow from them.

This brief video provides an excellent introduction to these core concepts from the perspective of an SRE.

How to get started with SLI/SLO with Steve McGhee

This video from "Is it Observable" features an ex-Google SRE, Steve McGhee, who explains the fundamentals of SLIs and SLOs.

Watch the initial segments that define what an SLI is and what an SLO is. Pay close attention to the idea of an SLI being a ratio and an SLO being a target for that ratio over time.

Choosing What to Measure: User Journeys and SLIs

The most effective SLIs are not arbitrary system metrics like CPU utilization. Instead, they are tied directly to critical user journeys. To define good SLIs, you must first ask: "What are the most important things a user does with my service, and what does 'working well' mean for each of those actions?"

As a team lead, you've likely had many discussions about prioritizing features. Thinking in terms of user journeys provides a structured way to ground those discussions in what actually matters to the customer.

Let's consider a typical e-commerce application. The image below breaks down a user journey and maps specific SLIs and SLOs to each step.

This diagram illustrates how different stages of an e-commerce user journey, from product search to order confirmation, can be measured with specific SLIs like latency, error rate, and availability, each with a corresponding SLO target.

As you can see, different interactions call for different types of indicators:

  • For "Searching for a product," latency is critical. A slow search is a bad experience.
  • For "Adding Product to Cart" and "Checkout," the error rate is paramount. Users must be able to perform these actions reliably.
  • For "Order Confirmation," availability is key. The user needs to know their order was successful.

The Google SRE workbook provides a systematic way to think about SLIs for different types of services.

Chapter 2 - Implementing SLOs - Google SRE

This chapter from Google's SRE Workbook is the definitive guide to implementing SLOs. This section provides a practical framework for getting started.

Please read the section A Worked Example. Focus on how it deconstructs an application into components (request-driven, pipeline, storage) and then suggests SLIs for each in Table 2-1. This structured approach is invaluable when you're analyzing a complex system.

The video below also provides a clear breakdown of the four most common SLI categories.

Service Level Objectives (SLOs) - noob to pro in under 30 minutes!

This video from Dark Mode Club explains the fundamental categories of SLIs.

Watch the segment from the book database, which covers the four basic categories of SLIs: latency, availability, consistency, and throughput.

Defining SLOs and the Power of the Error Budget

Once you have identified your SLIs, the next step is to define an SLO—the target. A common mistake is to aim for 100% reliability. While it sounds noble, it's an antipattern for several reasons:

  • Prohibitive Cost: Each additional "nine" of reliability (99%, 99.9%, 99.99%) comes at an exponentially increasing cost and complexity.
  • Diminishing Returns: The user's own environment (their device, their WiFi) is often far less reliable than your service. A 99.999% reliable service feels the same as a 99.9% reliable one if the user's internet connection drops.
  • Stifled Innovation: The biggest source of outages is change. If your goal is 100% reliability, you can never safely deploy new features or updates.

Instead of 100%, we choose a realistic target (e.g., 99.9%) that meets user expectations. The crucial insight is what this implies. The gap between your SLO and 100% is your Error Budget.

For a 99.9% availability SLO, you have a 0.1% error budget. This is the amount of unreliability you are allowed to have over the SLO's time window. The error budget is the central tool for data-driven decision-making in SRE.

  • If you have error budget remaining, you can safely deploy new features, perform maintenance, or run experiments that might introduce some risk.
  • If your error budget is exhausted, all non-essential development stops. The team's entire focus shifts to improving reliability and stability until the service is back within its SLO.

This creates a self-regulating feedback loop that balances innovation with reliability, resolving the classic conflict between "moving fast" and "not breaking things."

A Practical Guide to SLOs and SLIs in Microservices - sysctl.id

This article offers a very clear and concise explanation of SLOs and the error budget.

Read the section Defining SLOs. It perfectly summarizes why 100% is the wrong target and explains the strategic power of the error budget.

From Definition to Action

Defining SLIs and SLOs is just the first step. To make them useful, you need a process agreed upon by all stakeholders (Product, Development, and Operations/SRE).

The general workflow is as follows:

  1. Define SLIs and SLOs: Identify critical user journeys and set reliability targets, as we've discussed.
  2. Get Stakeholder Agreement: The product owner must agree the SLO is sufficient for user happiness. The development team must agree to abide by the error budget policy. The operations team must agree the SLO is achievable.
  3. Implement Measurement: Instrument your application code and infrastructure to export the metrics needed for your SLIs.
  4. Track and Alert: Use a monitoring system like Prometheus to track SLI performance against the SLO, calculate the remaining error budget, and alert when the budget is burning too quickly.
  5. Report and Review: Use dashboards to visualize SLO compliance and trends over time.

This image shows a typical SLO compliance report that might be shared with stakeholders. It provides a high-level view of service health, tracks trends, and flags services that are not meeting their objectives.

This scorecard shows quarterly SLO compliance for several services. Green indicates the service met its objectives, while red indicates a failure. The trends column shows whether reliability is improving or declining year-over-year (YoY) and quarter-over-quarter (QoQ).

For this to work, having a formal policy and clear documentation is non-negotiable.

Chapter 2 - Implementing SLOs - Google SRE

This part of the Google SRE workbook covers the crucial organizational aspects of making SLOs effective.

Please read the sections on Getting Stakeholder Agreement, Establishing an Error Budget Policy, and Documenting the SLO and Error Budget Policy. This is the "process" part of SRE that turns good ideas into effective practice.

Conclusion

In this lesson, we've established the language and framework for managing service reliability. You've moved beyond simple uptime metrics to a sophisticated, user-centric approach that is standard practice in high-performing engineering organizations.

Key Takeaways:

  • SLIs measure what matters: They are quantitative metrics tied directly to the user experience (e.g., availability, latency, error rate).
  • SLOs set the target: They define "good enough" reliability, acknowledging that 100% is an unrealistic and counterproductive goal.
  • The Error Budget drives decisions: This is your allowance for unreliability. It provides a data-driven framework for balancing feature development with stability work.
  • Reliability is a team sport: Effective SLOs require buy-in from product, development, and operations, formalized through documentation and an agreed-upon error budget policy.

Now that we understand what we need to measure and why, our next lesson will focus on the how. We will get hands-on and instrument a Go application to expose metrics for a Prometheus monitoring system, taking the first practical step toward tracking our SLIs.

Can't find a good explanation? Sign up and we'll make it for you

Sign up