Short answers to what readers ask most about this topic.
01What is the difference between SLI, SLO and SLA?
An SLI is a measurement of service quality, usually the ratio of good events to valid events. An SLO is the target you set for that SLI over a window, such as 99.9% of requests succeeding over 30 days. An SLA is a contract that attaches consequences, such as service credits, to missing a target.
02How do you calculate an error budget?
Subtract the SLO from 100% and multiply by the window length. For a 99.9% SLO over 30 days, 0.1% of 43,200 minutes is 43.2 minutes. You can do the same with requests: 0.1% of 1,000,000 valid requests is 1,000 failures.
03Why shouldn't an SLO be 100%?
Perfect reliability is unrealistic, and users cannot tell the difference because their own networks and devices fail more often. Each extra nine cuts the allowed failure tenfold and raises cost sharply. An SLO below 100% leaves an error budget that you can spend on releases and experiments.
04What happens when you burn through your error budget?
That is decided by your error budget policy, written before it happens. Typical actions are prioritising reliability work, freezing releases except urgent fixes until the service is back within its SLO, and holding a postmortem for large incidents. Without a policy, the SLO is only a report.
05How do you measure an availability SLI in Prometheus?
Divide the rate of failed requests by the rate of all requests from your http_requests_total counter, for example the sum of rate() over 5xx responses divided by the sum of rate() over all responses. The SLI is one minus that ratio. Record it per window as a recording rule so burn-rate alerts stay cheap.
An SLI is a measured ratio of good events to valid events, an SLO is the target you set for that ratio over a window, and an SLA is the contract with consequences if you miss it. Keep the SLO below 100% and stricter than the SLA. The gap is your error budget: 43.2 minutes per 30 days at 99.9%.
The three acronyms get used as synonyms in sprint planning, vendor contracts and status pages, and the confusion is expensive. A team promises 99.99% in a contract because the number sounded professional, then discovers its own deploys burn more than the allowed downtime.
This post defines each term with the wording from Google's SRE books, then does the arithmetic: how to write an SLI as a ratio, how to pick a target, how many minutes of failure each target allows, what to do when the budget runs out, how burn-rate alerts work, and the PromQL to measure it. The calculator output is from a run I did for this article, and every other figure is derived from those formulas or cited.
What is the difference between SLI, SLO and SLA?
The SRE book defines an SLI as a carefully defined quantitative measure of some aspect of the level of service provided, an SLO as a target value or range of values for a service level measured by an SLI, and an SLA as an explicit or implicit contract with users that includes consequences of meeting or missing the SLOs it contains. In short: measurement, goal, promise.
Term
Question it answers
Example
If it is missed
SLI (indicator)
How is the service doing right now?
Good HTTP requests divided by valid HTTP requests, measured over 5 minutes
Nothing. It is a measurement, not a commitment
SLO (objective)
What level are we aiming to deliver?
99.9% of valid requests succeed over a rolling 30 days
Internal consequences: the error budget policy applies
SLA (agreement)
What do we promise, and what do we owe if we break it?
99.5% monthly availability, otherwise a service credit (an illustrative figure)
Contractual: credits, refunds or termination rights
The SRE book gives a quick test for telling an SLO from an SLA: ask what happens if the target is not met. If there is no explicit consequence, you are looking at an SLO. Most internal products, including the ERP and POS backends I build, only ever need SLOs.
How do you write an SLI that reflects what users feel?
The SRE Workbook recommends expressing every SLI as the number of good events divided by the total number of valid events, so it runs from 0% (nothing works) to 100% (nothing is broken). A ratio is easy to alert on, easy to turn into a budget, and comparable across services. It also forces you to define two things: what counts as an event, and what counts as good.
Start from the user journey, not from the dashboard you already have. For a checkout screen in a POS, the journey is: the cashier taps pay, the sale is saved, the receipt data comes back. The Workbook groups typical SLIs by system type, and four of them cover most web backends:
Availability: the share of valid requests that return a non-error response. Count 5xx as bad; count 4xx as the caller's problem unless your own bug produces them.
Latency: the share of requests faster than a threshold, such as 300 ms. State it as a percentile or a threshold ratio, because the SRE book warns that a simple average can obscure tail latencies.
Correctness: the share of results that are right. For a pipeline or a ledger this can mean reconciliation checks that pass, not just requests that return 200.
Freshness: the share of data newer than a limit, for example stock levels synced from the branch within 5 minutes. It suits pipelines and caches, where a fast but stale answer is still a failure.
The Workbook suggests five or fewer SLI types covering the most critical functionality. More than that and nobody remembers which one is the real signal.
Separate the SLI specification from its implementation. The specification is the outcome you care about; the implementation is how you measure it. Load-balancer logs, server metrics and a synthetic probe are three implementations of the same availability SLI, each with different coverage and cost.
Why should an SLO be below 100%?
Because 100% is unreachable and chasing it is expensive. The SRE book calls it both unrealistic and undesirable to insist that SLOs be met 100% of the time: users sit behind networks, phones and ISPs that are less reliable than your service, so they cannot tell the difference, and the extra effort slows every release.
Each additional nine cuts the allowed failure by a factor of ten. Going from 99.9% to 99.99% shrinks a 30-day budget from 43.2 minutes to 4.32 minutes, so a single 5-minute bad deploy already breaks the target. The cost is real redundancy, rollout machinery and on-call rotations. My post on high availability, redundancy and failover design covers what those nines cost in infrastructure; here the point is only that the target has to be one you can afford.
Do not copy the target from today's performance either. The SRE book advises against picking a target based on current performance, because it locks in whatever accident you happened to measure. Choose it from what users need, then check whether you can meet it. For a POS API on a single VPS, I would start at 99.5% and earn my way up, not announce four nines and hope.
How does an SLA relate to an SLO?
The SLA is the contractual consequence of an SLO, and the SLO behind it should be stricter than what the contract promises. The SRE book recommends a tighter internal SLO than the one advertised to users, which gives you room to fix chronic problems before customers notice and room to trade performance for cost.
Using the illustrative figures from the table: an SLA of 99.5% allows 216 minutes of failure per 30 days (0.005 times 43,200 minutes), while an internal SLO of 99.9% allows 43.2. The internal alarm rings long before a credit is owed. If the SLO and the SLA were the same number, your first warning would be the invoice.
How do you calculate an error budget?
The error budget is 100% minus the SLO. Multiply it by the window length to get time, or by the number of valid requests to get a count of failures you can afford. A 30-day window has 30 times 24 times 60 = 43,200 minutes, so each SLO converts directly.
SLO
Budget per 30 days
Budget in seconds
Failed requests per 1M
99%
432 minutes (7.2 hours)
25,920
10,000
99.9%
43.2 minutes
2,592
1,000
99.95%
21.6 minutes
1,296
500
99.99%
4.32 minutes
259.2
100
Rather than trust my mental arithmetic, I wrote the calculator in TypeScript and ran it. It prints the budget for each SLO and, for the burn-rate section below, how fast a given burn rate spends a 99.9% budget.
const WINDOW_DAYS = 30;
const WINDOW_MINUTES = WINDOW_DAYS * 24 * 60; // 43200
const WINDOW_HOURS = WINDOW_DAYS * 24; // 720
// The budget is whatever the SLO leaves over: 100% minus the target.
const budgetFraction = (slo: number): number => 1 - slo;
const budgetMinutes = (slo: number): number =>
budgetFraction(slo) * WINDOW_MINUTES;
// With 1,000,000 valid requests in the window, how many may fail?
const budgetRequests = (slo: number, validRequests: number): number =>
budgetFraction(slo) * validRequests;
// Burn rate 1 spends the budget in exactly the window. Burn rate B spends it in window / B.
const hoursToExhaust = (burnRate: number): number => WINDOW_HOURS / burnRate;
// Share of the whole window's budget consumed by one alert window at a given burn rate.
const budgetConsumed = (burnRate: number, alertWindowHours: number): number =>
(burnRate * alertWindowHours) / WINDOW_HOURS;
for (const slo of [0.99, 0.999, 0.9995, 0.9999]) {
const mins = budgetMinutes(slo);
console.log(
(slo * 100).toFixed(2).padStart(6) + "% " +
mins.toFixed(2).padStart(7) + " min " +
(mins * 60).toFixed(1).padStart(8) + " s " +
budgetRequests(slo, 1_000_000).toFixed(0).padStart(6) + " failed of 1M requests",
);
}
// [burn rate, long window in hours, short window in minutes]
const rules: Array<[number, number, number]> = [
[14.4, 1, 5],
[6, 6, 30],
[1, 72, 360],
];
for (const [burn, longHours, shortMinutes] of rules) {
console.log(
"burn " + String(burn).padEnd(4) +
" long " + String(longHours).padStart(2) + " h, short " + String(shortMinutes).padStart(3) + " min" +
" error ratio above " + (burn * budgetFraction(0.999)).toFixed(4) +
" consumes " + (budgetConsumed(burn, longHours) * 100).toFixed(1) + "%" +
" budget gone in " + hoursToExhaust(burn).toFixed(1) + " h",
);
}
This is the real output from running it with Node.js type stripping:
$ node budget-calc.ts
99.00% 432.00 min 25920.0 s 10000 failed of 1M requests
99.90% 43.20 min 2592.0 s 1000 failed of 1M requests
99.95% 21.60 min 1296.0 s 500 failed of 1M requests
99.99% 4.32 min 259.2 s 100 failed of 1M requests
burn 14.4 long 1 h, short 5 min error ratio above 0.0144 consumes 2.0% budget gone in 50.0 h
burn 6 long 6 h, short 30 min error ratio above 0.0060 consumes 5.0% budget gone in 120.0 h
burn 1 long 72 h, short 360 min error ratio above 0.0010 consumes 10.0% budget gone in 720.0 h
The numbers match the table. They also match the 43.2 minutes quoted for 99.9% over 30 days. Time-based budgets assume traffic is steady; a request-based budget (the last column) is fairer, because a ten-minute outage at 3 a.m. costs fewer failed requests than the same outage at the lunch peak.
The window is part of the SLO. 99.9% over a day allows 86.4 seconds, 99.9% over a quarter allows 129.6 minutes (0.001 times 129,600). Say which window you mean, and prefer a rolling window to a calendar month so the budget does not reset on the first of the month.
What happens when the error budget is spent?
An error budget policy decides, in advance, what the team does when the budget is gone. The Workbook is blunt that without a policy, SLO compliance is just another reporting metric. Write it down, name an owner and an escalation path, and include actions such as these:
Prioritise reliability bugs and postmortem action items over new features.
Freeze releases until the service is back within its SLO. The Workbook's example halts all changes other than P0 issues and security fixes.
Require a postmortem when one incident consumes a large share of the budget. The example uses more than 20% of the budget over four weeks.
Decide how outages caused by a dependency count. The Workbook favours a freeze regardless of cause, but says the right choice depends on the service and belongs in the policy.
The reverse also matters. A budget that is barely touched is a signal to ship faster or take more risk, because reliability you do not need is velocity you gave away.
How do burn-rate alerts work?
Alerting on every error spike is noise, and alerting on a raw threshold ignores the budget. The Workbook defines burn rate as how fast, relative to the SLO, the service consumes the error budget. Burn rate 1 spends the budget in exactly the SLO window; burn rate 14.4 spends it 14.4 times faster.
The recommended multi-window, multi-burn-rate setup for a 99.9% SLO pages on fast burns and opens tickets on slow ones. Each alert needs two windows to be true at once: a long window to show the burn is significant, and a short window, about one twelfth of the long one, to show it is still happening. The last column is derived from burn rate and the 720-hour window.
Severity and burn rate
Long and short window
Budget consumed
Budget gone in
Page at 14.4
1 hour and 5 minutes
2%
50 hours
Page at 6
6 hours and 30 minutes
5%
120 hours
Ticket at 1
3 days and 6 hours
10%
720 hours
The arithmetic is short: 14.4 times 1 hour divided by 720 hours is 2%, which is why a one-hour burn at that rate is worth waking someone up. The short window also stops an alert from staying red for an hour after you have already fixed the problem.
How do you measure an availability SLI in Prometheus?
If your service exposes an http_requests_total counter, the SLI is a ratio of rates. The Prometheus docs say rate() takes a range vector, should only be used with counters, and adjusts for counter resets after restarts. Record the error ratio once per window so alerts stay cheap. The label that holds the status code depends on your client library; I use code here, and many NestJS setups name it status_code.
groups:
- name: api-slo
rules:
# Bad events / valid events over 5 minutes. 4xx is the caller's mistake, so it
# stays in the denominator but never in the numerator.
- record: job:slo_errors_per_request:ratio_rate5m
expr: |
sum by (job) (rate(http_requests_total{job="api", code=~"5.."}[5m]))
/
sum by (job) (rate(http_requests_total{job="api"}[5m]))
# Same expression, longer range. Repeat for 30m, 6h and 3d.
- record: job:slo_errors_per_request:ratio_rate1h
expr: |
sum by (job) (rate(http_requests_total{job="api", code=~"5.."}[1h]))
/
sum by (job) (rate(http_requests_total{job="api"}[1h]))
The same counters give the SLI over the whole SLO window. For latency, a histogram lets you divide the requests under a threshold by all requests, which is exact, unlike an interpolated percentile:
# Availability SLI over the SLO window: good / valid. Run it as a recording rule in
# production, because a 30d range query over raw counters is slow.
1 - (
sum(increase(http_requests_total{job="api", code=~"5.."}[30d]))
/
sum(increase(http_requests_total{job="api"}[30d]))
)
# Latency SLI: share of requests faster than 300 ms. The le="0.3" bucket must exist
# in your histogram, so pick the threshold from your bucket boundaries, not the reverse.
sum(rate(http_request_duration_seconds_bucket{job="api", le="0.3"}[5m]))
/
sum(rate(http_request_duration_seconds_count{job="api"}[5m]))
# p99 for dashboards (keep le in the aggregation). Do not alert on this one.
histogram_quantile(0.99,
sum by (le) (rate(http_request_duration_seconds_bucket{job="api"}[5m])))
Then wire the alerts from the table. These expressions follow the Workbook's examples for a 99.9% SLO, so the budget fraction is 0.001:
groups:
- name: api-slo-alerts
rules:
# SLO 99.9%, so the budget fraction is 0.001. Both windows must exceed the
# threshold: the long one proves it is significant, the short one proves it is
# still happening.
- alert: ApiErrorBudgetFastBurn
expr: |
(
job:slo_errors_per_request:ratio_rate1h{job="api"} > (14.4 * 0.001)
and
job:slo_errors_per_request:ratio_rate5m{job="api"} > (14.4 * 0.001)
)
or
(
job:slo_errors_per_request:ratio_rate6h{job="api"} > (6 * 0.001)
and
job:slo_errors_per_request:ratio_rate30m{job="api"} > (6 * 0.001)
)
labels:
severity: page
- alert: ApiErrorBudgetSlowBurn
expr: |
job:slo_errors_per_request:ratio_rate3d{job="api"} > 0.001
and
job:slo_errors_per_request:ratio_rate6h{job="api"} > 0.001
labels:
severity: ticket
A service with no traffic yields 0 divided by 0, which Prometheus returns as NaN, so the alert silently never fires. Add a synthetic probe or an absent() alert so that silence on a counter is itself noticed.
The rule I carry away: an SLI is a ratio you can compute, an SLO is a number below 100% you can afford, and an SLA is only what you are willing to pay for missing it, set looser than the SLO. Start with one availability SLI and one latency SLI, convert the target to minutes, write the policy before the first outage, and alert on burn rate instead of raw errors.