You have mastered the three pillars of observability: metrics with Prometheus, logs with Loki and the ELK Stack, and distributed tracing with OpenTelemetry. Data alone does not make a system reliable. This lesson turns that data into decisions: you will define Service Level Indicators (SLIs), set Service Level Objectives (SLOs), and use error budgets to balance feature velocity with reliability - the practice that separates a monitored system from an SRE-managed one.

1. Learning Objectives

By the end of this lesson, you will be able to:

  • Explain the difference between an SLI, an SLO, and an error budget
  • Choose meaningful SLIs for availability, latency, and correctness
  • Set SLO targets that reflect user expectations instead of aspirational 100% values
  • Calculate a monthly error budget from an SLO target in seconds
  • Measure an availability SLI with Prometheus queries over a rolling window
  • Alert on error budget burn rate and explain when to halt deploys

2. Why This Matters

Imagine you run an e-commerce platform that earns revenue on every request. Your monitoring stack records every metric, log, and trace, but nobody can answer the simplest question: is the system reliable enough for our users? Without a target, a 99.5% uptime month (about 3.6 hours of downtime) either passes silently or triggers a fire drill, depending on who happens to be on call.

SLOs answer that question with numbers. They convert raw observability data into a contract between the engineering team and the business: this service will be available 99.9% of the time over 30 days, and here is how we measure it. Error budgets then turn that contract into a decision tool. When the budget is full, you can ship features with confidence. When it is nearly exhausted, you stop risky changes and focus on reliability work.

This lesson completes the observability arc you started with Prometheus, Loki, and OpenTelemetry. Those tools tell you what happened; SLOs tell you whether it matters and what to do about it.

3. Core Concepts

What is an SLI?

A Service Level Indicator (SLI) is a quantitative measurement of a specific aspect of your service, such as the fraction of requests that succeed or the proportion of time the service is reachable. A good SLI is measurable from existing telemetry, meaningful to users, and stable over time. The classic pair for any request-driven service is availability (successful requests divided by total requests) and latency (how fast requests complete).

# sli.yaml - declarative SLI/SLO definition for the api-gateway service
service: api-gateway
slis:
  - name: availability
    indicator: successful requests / total requests
    source: prometheus
    query: 'sum(rate(http_requests_total{status_code!~"5.."}[5m])) / sum(rate(http_requests_total[5m]))'
  - name: latency_p95
    indicator: 95th percentile of request duration
    source: prometheus
    query: 'histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (le))'
slos:
  - name: api-availability-30d
    sli: availability
    target: 0.999
    window: 30d
  - name: api-latency-30d
    sli: latency_p95
    target: 0.300
    unit: seconds
    window: 30d
SLI definitions for a typical API service

What is an SLO?

A Service Level Objective (SLO) is the target you set for an SLI, usually over a fixed window such as 30 days. For example, availability of at least 99.9% over 30 days, or p95 latency under 300 milliseconds. The SLO is not a guess: it should reflect what users actually need and what the system can realistically deliver. Each SLO needs an owner, a measurement source, and a review process, because targets that nobody owns are targets that never change.

What is an Error Budget?

The error budget is the amount of unreliability your SLO permits: 100% minus the SLO target. For a 99.9% availability SLO over 30 days, the budget is 0.1%, or about 2,592 seconds (43 minutes) of allowed errors per month. Every 5xx response, timeout, or outage spends from that budget. When the budget is spent, the team's job shifts from shipping features to restoring reliability, because further failures would break the user-facing contract.

SLI Types and Examples

Different services need different SLIs. Start with the ones that reflect user-visible behavior:

  • Availability - the fraction of requests that succeed or the proportion of time the service is up. The default choice for most services.
  • Latency - response time percentiles such as p50, p95, or p99. Choose the percentile that matches the user experience you care about.
  • Throughput - requests processed per second. Important for capacity planning and batch systems.
  • Durability - the proportion of writes that are durably stored. Critical for databases and message queues.
  • Correctness - the fraction of responses that are not just fast, but right. Used by search, recommendation, and data pipelines.

Choosing SLO Targets

The most common mistake is setting targets you cannot defend. A 100% availability target leaves zero error budget, so every tiny blip becomes a page and the team can never ship. Industry practice suggests starting near 99.9% for user-facing services and 99.99% for critical data planes, then adjusting based on what users actually experience. Each additional nine roughly multiplies the operational cost, so pick the loosest target your users accept, not the strictest one you can imagine.

4. Hands-On Practice

In this practice you will define SLIs and SLOs for a sample api-gateway service, measure the availability SLI with Prometheus, calculate the error budget in Python, and wire up burn-rate alerting. You need a Prometheus instance scraping the service and its HTTP metrics.

Step 1: Define SLIs and SLOs for the api-gateway Service

Start by writing a declarative definition of your indicators and targets. Keeping this in a file makes the contract explicit and reviewable.

# sli.yaml - declarative SLI/SLO definition for the api-gateway service
service: api-gateway
slis:
  - name: availability
    indicator: successful requests / total requests
    source: prometheus
    query: 'sum(rate(http_requests_total{status_code!~"5.."}[5m])) / sum(rate(http_requests_total[5m]))'
  - name: latency_p95
    indicator: 95th percentile of request duration
    source: prometheus
    query: 'histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (le))'
slos:
  - name: api-availability-30d
    sli: availability
    target: 0.999
    window: 30d
  - name: api-latency-30d
    sli: latency_p95
    target: 0.300
    unit: seconds
    window: 30d
sli.yaml - SLI and SLO definitions

Step 2: Measure the Availability SLI in Prometheus

The api-gateway exposes http_requests_total with a status_code label. The availability SLI over the last 30 days is the ratio of non-5xx responses to all responses. The latency SLI comes from a histogram, and Blackbox probes give you an independent availability reading from the user's point of view.

# Availability SLI: fraction of requests that succeeded over 30 days
sum(rate(http_requests_total{status_code!~"5.."}[30d]))
/
sum(rate(http_requests_total[30d]))

# Latency SLI: p95 response time over the last 5 minutes
histogram_quantile(
  0.95,
  sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)

# Availability from Blackbox probes (0/1 gauge averaged over 30 days)
avg_over_time(probe_success{instance="https://api.example.com"}[30d])
PromQL queries for availability and latency SLIs

Step 3: Calculate the Error Budget with Python

Write a small calculator that converts an SLO percentage into a concrete monthly budget and tracks how much of it remains. This is the number your team should see on every dashboard.

#!/usr/bin/env python3
# error_budget.py - monthly error budget calculator

MONTHLY_SECONDS = 30 * 24 * 60 * 60  # 2,592,000 seconds

def slo_budget(slo_percent, window_seconds=MONTHLY_SECONDS):
    allowed_unavailable = window_seconds * (1 - slo_percent / 100.0)
    return {
        'slo': slo_percent,
        'window_seconds': window_seconds,
        'allowed_errors_seconds': round(allowed_unavailable, 1),
        'allowed_errors_percent': round(100 - slo_percent, 4),
    }

def error_budget_remaining(slo_percent, actual_percent):
    # Fraction of budget left: 1.0 = full budget, 0.0 = exhausted
    budget = 100 - slo_percent
    spent = 100 - actual_percent
    return max(0.0, min(1.0, 1 - spent / budget))

if __name__ == '__main__':
    for slo in (99.9, 99.95, 99.99):
        print(slo_budget(slo))
    print('remaining:', error_budget_remaining(99.9, 99.82))
error_budget.py - monthly budget calculator
kubectl get pods -w
$ python3 error_budget.py
{'slo': 99.9, 'window_seconds': 2592000, 'allowed_errors_seconds': 2592.0, 'allowed_errors_percent': 0.1}
{'slo': 99.95, 'window_seconds': 2592000, 'allowed_errors_seconds': 1296.0, 'allowed_errors_percent': 0.05}
{'slo': 99.99, 'window_seconds': 2592000, 'allowed_errors_seconds': 259.2, 'allowed_errors_percent': 0.01}
remaining: 0.2
Running the error budget calculator

The output shows that a 99.9% SLO allows 2,592 seconds (about 43 minutes) of errors per 30-day window. With actual availability of 99.82%, the remaining budget is 0.2 (20% left), which means the month is not going well and the team should slow down feature work before the budget is exhausted.

Step 4: Alert on Error Budget Burn Rate

Alerting on the raw error rate pages your team on every blip. Instead, alert on burn rate: how many times faster the budget is being consumed than allowed. A burn rate of 1 means you will exhaust the budget exactly at the end of the window; 14.4 means the monthly budget would be gone in about two days. The classic multi-window pattern uses a fast window (1 hour) to catch problems immediately and a slow window (6 hours) to confirm they persist.

# Prometheus alerting rule: error budget burn rate
groups:
  - name: slo-burn-rate
    rules:
      # Burn rate = observed error rate / allowed error rate (1 - SLO)
      # For a 99.9% SLO the allowed error rate is 0.001 (0.1%).
      # A burn rate of 14.4 consumes about 2% of the monthly budget in 1 hour.
      - alert: ApiErrorBudgetBurnRateHigh
        expr: |
          (
            sum(rate(http_requests_total{status_code=~"5.."}[1h]))
            /
            sum(rate(http_requests_total[1h]))
          ) / 0.001 > 14.4
        for: 15m
        labels:
          severity: page
          team: platform
        annotations:
          summary: 'API error budget burning at {{ $value }}x'
          description: 'Error budget consumed 14.4x faster than allowed in the last hour.'
Prometheus alerting rule for burn rate

Step 5: Track the Budget in Grafana

Add a panel to your Grafana dashboard that shows the remaining error budget as a fraction from 1.0 down to 0.0, and a second panel with the current burn rate. When the budget line approaches zero, the panel turns red and the team knows to stop shipping risky changes.

# Error budget remaining as a fraction (1.0 = full, 0.0 = exhausted)
# for a 99.9% SLO over a 30-day window
1 - (
  sum(rate(http_requests_total{status_code=~"5.."}[30d]))
  /
  sum(rate(http_requests_total[30d]))
  / 0.001
)

# Burn rate over the last hour (panel or alert input)
sum(rate(http_requests_total{status_code=~"5.."}[1h]))
/
sum(rate(http_requests_total[1h]))
/ 0.001
Grafana panel queries for budget remaining and burn rate

5. Building an SLO Policy

An SLO without a policy is just a number. Real teams write down what happens when the budget is spent, so the response is automatic instead of emotional. Here is a policy that works for a platform team running an api-gateway:

  • 99.9% availability over 30 days for the public API.
  • p95 latency under 300 ms over 30 days for all traffic.
  • Error budget for each SLO is 0.1% of the window (2,592 seconds per month).
  • Deploys are frozen when either budget drops below 30% remaining.
  • All pages during an error budget deficit require a written postmortem.
  • SLOs are reviewed every quarter with product and engineering leadership.

Notice what the policy does not say: it does not say aim for 100% availability, and it does not promise a specific number of nines for internal services. The policy ties reliability decisions to the budget, so a team that ships a risky feature after the budget is exhausted has violated an explicit, agreed rule rather than an implicit expectation.

Ownership matters too. Assign one engineer to each SLO as its owner. The owner watches the burn rate, proposes target changes, and leads the response when the budget is in danger. Ownership turns a dashboard widget into a living operational contract.

6. Common Errors and Solutions

1. Confusing the SLI with the SLO. The SLI is the measurement (the fraction of successful requests); the SLO is the target on that measurement (at least 99.9% over 30 days). Teams that mix them up write alerts against raw metrics with no window, producing noise and no decisions. Always state all three: indicator, target, and window.

2. Setting a 100% availability target. A perfect target leaves an empty error budget, so the smallest blip exhausts it and every incident becomes an emergency. Real systems have planned maintenance, deploys, and upstream failures. Pick the loosest target users accept, typically 99.9% to 99.99%, and keep a small buffer for the unexpected.

3. Measuring the wrong request population. Counting internal health checks or retries inflates the SLI, while ignoring timeouts undercounts failures. The SLI must match the user experience: exclude synthetic probes from the availability ratio, count timeouts as errors, and never include 4xx client errors in the failure ratio, because they are not service failures at all.

4. Alerting on every SLO breach instead of burn rate. A breach alert fires only after the budget is already spent, when it is too late to act. Burn-rate alerts fire while there is still budget left to protect. For a 99.9% SLO, alert when the hourly burn rate exceeds 14.4x for 15 minutes, which signals roughly 2% of the monthly budget consumed in one hour.

5. Choosing a window that is too short. A 24-hour SLO window makes weekend traffic patterns exhaust the budget randomly, while a 90-day window hides real degradation for months. The 30-day window is the industry default: long enough to absorb routine noise, short enough to keep the budget meaningful. Review the window whenever the service's traffic profile changes.

7. Summary Checklist

  • Defined SLIs for availability, latency, and at least one business-critical behavior
  • Set SLO targets with windows (for example, 99.9% over 30 days) instead of bare percentages
  • Calculated the monthly error budget in seconds and displayed it on a dashboard
  • Wrote Prometheus queries that measure each SLI from existing metrics
  • Added burn-rate alerts with fast and slow windows instead of raw threshold alerts
  • Documented an error budget policy that defines when deploys freeze
  • Assigned an owner to every SLO and scheduled a quarterly review

8. Practice Exercise

Challenge: build an SLO for your checkout service. Choose one user-facing service you operate (or the sample api-gateway) and complete the following:

  • Define two SLIs: availability and p95 latency, with the exact Prometheus queries that measure them.
  • Set SLO targets and windows based on what your users actually need, and justify each target in one sentence.
  • Calculate the error budget in seconds for a 30-day window.
  • Write a burn-rate alerting rule with fast (1h) and slow (6h) windows, and state which burn-rate values trigger page versus ticket.
  • Sketch a Grafana panel showing budget remaining, and define the freeze rule when it drops below 30%.

Acceptance criteria: a reviewer can read your SLO document and answer, without asking you: what is measured, what is the target, how much budget is left today, and what happens when the budget is exhausted.

9. Next Steps

You have completed the Monitoring and Observability series: system visibility, Prometheus metrics, Grafana dashboards, centralized logging, alerting and incident response, distributed tracing, and now SLOs with error budgets. This concludes the Monitoring and Observability category. Next, the rotation moves to the DevOps and SRE series, where these reliability targets become the foundation for capacity planning and the day-to-day practice of SRE teams.

Gataya Med

DevOps Engineer & Backend Developer. Sharing insights on cloud, automation, and scalable systems.

Comments (0)

Sarah Chen August 12, 2026

This is exactly what I needed! The initContainer approach solved our migration issues completely. Thanks for the detailed guide!

Reply

Leave a Comment