Short answers to what readers ask most about this topic.
01How do you design a metrics and monitoring system?
Design it as a pipeline of six stages: instrumentation, collection, time-series storage, query, alerting and dashboards. Decide the label set and scrape interval first because they set storage cost, then choose retention and write alerts on user-facing symptoms. Estimate series times samples per second times bytes per sample before you ship any new metric.
02What is the difference between pull and push monitoring?
In pull monitoring the collector visits each target's metrics endpoint on a schedule, so a failed scrape tells you the target is down. In push monitoring the application sends its metrics to a collector, which suits short-lived batch jobs but makes silence ambiguous. Prometheus pulls by default and uses a Pushgateway for batch jobs.
03What is cardinality explosion in Prometheus?
Every unique combination of label values is a separate time series, so series count is the product of each label's distinct values. One latency histogram with 5 methods, 40 routes, 6 status codes and 10 instances already makes 168,000 series. Adding a label with 10,000 values such as a user ID multiplies that into the billions.
04How much storage does a Prometheus server need?
Disk space equals retention in seconds times ingested samples per second times bytes per sample, and the Prometheus docs say samples average around 1 to 2 bytes. As an assumption, 40,000 series scraped every 15 seconds at 2 bytes per sample needs about 6.4 GiB for 15 days. Doubling the scrape interval halves that figure.
05How do I avoid alert fatigue?
Page only on symptoms users feel, such as error ratio or latency, and put causes like CPU on dashboards. Add a for duration so brief blips never fire, and delete any alert where the on-call person has no action to take. Fewer, actionable alerts are trusted, and trusted alerts get answered.
Design a Metrics and Monitoring System: A Prometheus Pipeline
How to design a metrics and monitoring system end to end: instrumentation, pull versus push, time-series storage maths, cardinality, retention and alerting.
To design a metrics and monitoring system, build a pipeline: instrument services with counters, gauges and histograms, collect by scraping, store in a time-series database, query, then alert on user-facing symptoms. Size storage as series times samples per second times roughly two bytes, and cap label cardinality before it multiplies.
On a single VPS running a NestJS API, Postgres and Redis in Docker, the first monitoring decision is rarely which tool to install. Prometheus is the default answer. What decides whether the setup stays useful is everything around the install: what you measure, how many label combinations you allow, how long you keep it, and what is allowed to wake someone up.
This is the whole-system view. It walks the pipeline from instrumentation to dashboards, then spends its time on the three places designs go wrong: storage arithmetic, cardinality and alert design. Every figure comes from the Prometheus documentation or from an estimator script whose inputs are labelled as assumptions, and whose output below is a real run.
What does a metrics and monitoring system consist of?
Six stages, each with one job and one characteristic way of failing. Naming them separately matters because the fix for a noisy pager (stage five) is never in the stage where people look first (stage six).
Stage
Its job
How it usually fails
1. Instrumentation
Code exposes counters, gauges and histograms with labels
A label with unbounded values creates unbounded series
2. Collection
A scraper pulls (or clients push) samples every interval
Missed scrapes leave gaps; short-lived jobs are never seen
3. Storage
A time-series database keeps series, timestamps and values
Disk grows with series count times retention
4. Query
A query language computes rates, ratios and quantiles
Long-range queries over raw data are slow
5. Alerting
Rules are evaluated on a schedule and routed to people
Cause-based alerts with no duration cause alert fatigue
6. Dashboards
Graphs for diagnosis once an alert has fired
Forty panels nobody opens; no one knows where to look first
Google's SRE book frames the goal of this whole pipeline as answering two questions: what is broken, and why. The pipeline gives you the data for the second question, but only the alerting layer is responsible for the first, which is why the rest of this post keeps returning to it.
Which metric types does a monitoring system need?
The Prometheus documentation defines four core types. Three cover almost everything an application needs.
Counter: a cumulative value that only goes up, or resets to zero on restart. Use it for requests, errors and bytes sent, and always query it with rate(), never the raw number.
Gauge: a value that can rise and fall, such as queue depth, memory in use or open connections. You read it directly.
Histogram: counts observations into configurable buckets and also exposes a sum and a count. It is the right type for latency because quantiles can be calculated from it and aggregated across instances.
The fourth type, the summary, calculates quantiles on the client, and those cannot be meaningfully averaged across instances. For a fleet of API replicas that is usually the reason to choose a histogram. Here is the instrumentation shape with prom-client. Note that the route label is the route template, not the raw URL.
import { Counter, Histogram, collectDefaultMetrics, register } from "prom-client";
collectDefaultMetrics(); // process and event-loop gauges for free
// Counter: only goes up. Query it with rate(), never read the raw value.
const requests = new Counter({
name: "http_requests_total",
help: "Requests handled",
labelNames: ["method", "route", "status"], // route is the TEMPLATE "/orders/:id", never the raw URL
});
// Histogram: cumulative buckets, so quantiles can be aggregated across instances later.
const latency = new Histogram({
name: "http_request_duration_seconds",
help: "Request latency",
labelNames: ["route"],
buckets: [0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10], // 11 finite buckets
});
// In the handler:
// const end = latency.startTimer({ route: "/orders/:id" });
// ...
// end(); requests.inc({ method: "GET", route: "/orders/:id", status: "200" });
// The scrape endpoint Prometheus pulls from:
// app.get("/metrics", async (_, res) => res.type(register.contentType).send(await register.metrics()));
Should metrics be pulled or pushed?
Prometheus pulls: it visits each target's metrics endpoint on an interval. Its own FAQ says pulling is slightly better than pushing, while accepting that push has its place. The trade-offs are concrete.
Concern
Pull (scrape)
Push
Is the target alive?
A failed scrape is recorded as up equal to 0, so death is itself a signal
Silence is ambiguous: dead, or just quiet?
Short-lived batch jobs
May finish before the next scrape and never be seen
Natural fit: push the result once at the end
Network shape
Scraper must reach every target
Targets need only outbound access to the collector
Rate control
The collector decides the interval centrally
A misbehaving client can flood the collector
My default for long-running services is pull, with a Pushgateway only for batch jobs such as a nightly report or backup. The config below shows both: two API instances scraped directly, a label dropped at ingestion, and a Pushgateway scraped like any other target.
global:
scrape_interval: 15s # every series costs one sample per interval, so this is a storage knob
evaluation_interval: 15s # how often alert and recording rules run
rule_files:
- /etc/prometheus/rules/*.yml
scrape_configs:
# Pull: Prometheus visits each target's /metrics endpoint.
- job_name: api
metrics_path: /metrics
static_configs:
- targets: ["api-1:3000", "api-2:3000"]
metric_relabel_configs:
# Drop a label you should never have shipped before it becomes a series.
- action: labeldrop
regex: user_id
# Short-lived jobs cannot be scraped, so they push to a Pushgateway that Prometheus scrapes.
- job_name: pushgateway
honor_labels: true # keep the job label the batch job set, not "pushgateway"
static_configs:
- targets: ["pushgateway:9091"]
- job_name: node
static_configs:
- targets: ["node-exporter:9100"]
How much storage does a metrics system need?
The Prometheus storage documentation gives the formula: needed disk space equals retention time in seconds, times ingested samples per second, times bytes per sample. It states that Prometheus uses on average only around 1 to 2 bytes per sample. I use 2 as a conservative planning figure. Samples per second is simply series divided by the scrape interval. The script below applies that formula, and the inputs are assumptions I chose, not measurements.
// Cardinality and storage estimator. Every input is an assumption, not a measurement.
const BYTES_PER_SAMPLE = 2; // upper end of the "1-2 bytes per sample" in the Prometheus storage docs
const SECONDS_PER_DAY = 86_400;
// Series for one metric = product of label value counts, times series per label set.
const product = (counts: number[]): number => counts.reduce((a, b) => a * b, 1);
// A histogram with 11 finite buckets exposes 11 + 1 (+Inf) bucket series, plus _sum and _count.
const HISTOGRAM_SERIES_PER_LABEL_SET = 11 + 1 + 2;
const labels = { method: 5, route: 40, status: 6, instance: 10 };
const labelSets = product(Object.values(labels));
const histogramSeries = labelSets * HISTOGRAM_SERIES_PER_LABEL_SET;
console.log("label sets (method x route x status x instance):", labelSets);
console.log("histogram series:", histogramSeries);
const withUserId = histogramSeries * 10_000;
console.log("same histogram plus a user_id label (10,000 users):", withUserId);
const storage = (series: number, scrapeSeconds: number, retentionDays: number) => {
const samplesPerSecond = series / scrapeSeconds;
const bytesPerDay = samplesPerSecond * SECONDS_PER_DAY * BYTES_PER_SAMPLE;
const totalGiB = (bytesPerDay * retentionDays) / 1024 ** 3;
return { samplesPerSecond, mbPerDay: bytesPerDay / 1024 ** 2, totalGiB };
};
const scenarios: Array<[string, number, number, number]> = [
["40,000 series, 15 s, 15 d", 40_000, 15, 15],
["40,000 series, 60 s, 15 d", 40_000, 60, 15],
["40,000 series, 15 s, 365 d raw", 40_000, 15, 365],
["histogram + user_id, 15 s, 15 d", withUserId, 15, 15],
];
for (const [name, series, interval, days] of scenarios) {
const r = storage(series, interval, days);
console.log(name.padEnd(34), r.samplesPerSecond.toFixed(0).padStart(10), "samples/s",
r.mbPerDay.toFixed(1).padStart(12), "MiB/day", r.totalGiB.toFixed(1).padStart(10), "GiB");
}
Run with Node, this is the real output. Forty thousand series is an assumed fleet size, and the histogram line is the cardinality example from the next section.
$ node metrics-estimator.ts
label sets (method x route x status x instance): 12000
histogram series: 168000
same histogram plus a user_id label (10,000 users): 1680000000
40,000 series, 15 s, 15 d 2667 samples/s 439.5 MiB/day 6.4 GiB
40,000 series, 60 s, 15 d 667 samples/s 109.9 MiB/day 1.6 GiB
40,000 series, 15 s, 365 d raw 2667 samples/s 439.5 MiB/day 156.6 GiB
histogram + user_id, 15 s, 15 d 112000000 samples/s 18457031.3 MiB/day 270366.7 GiB
Three things fall out of the arithmetic. First, 40,000 series at a 15 second interval costs about 6.4 GiB for 15 days, which fits easily on a small VPS disk. Second, moving to a 60 second interval cuts that to a quarter, because samples per second is inversely proportional to the interval. Third, keeping the same raw data for a year costs about 157 GiB, which is why retention is a design decision and not a default. The last line is the warning: one badly labelled histogram would need hundreds of terabytes.
What is cardinality explosion and how do you prevent it?
The Prometheus naming guidance is blunt: every unique combination of label key and value pairs is a new time series, and you should not put high-cardinality values such as user IDs or email addresses in labels. The arithmetic shows why. Series count is the product of every label's value count, multiplied by the series each label set produces.
In the estimator, a latency histogram with 5 methods, 40 routes, 6 status codes and 10 instances has 12,000 label sets. Each produces 14 series (11 finite buckets, one +Inf bucket, a sum and a count), so one metric is 168,000 series, already more than the whole 40,000 series fleet budget. Adding a user_id label with 10,000 values multiplies that to 1.68 billion. Four guardrails keep this from happening.
Label only with bounded sets: HTTP method, route template, status class. Never raw URLs, user IDs, order numbers or free text.
Estimate before you ship: multiply the label value counts on the pull request that adds a metric.
Drop dangerous labels at ingestion with metric_relabel_configs, as the scrape config above does.
Alert on the Prometheus server's own series count so a cardinality jump is noticed in hours, not at the disk-full page.
Cardinality is multiplicative, not additive. Adding one label with N values does not add N series, it multiplies the existing count by N. The most common cause is a route label filled with the raw path, so /orders/1042 and /orders/1043 become different series.
How long should you keep metrics, and what is downsampling?
Retention is a storage cost you choose. Prometheus keeps 15 days by default and exposes a retention time flag to change it, and the same documentation lets you cap by size instead. Most operational questions (what changed in the last deploy, is this slower than yesterday) need days, not years. Keeping raw 15 second data for a year, as the estimator showed, mostly buys disk bills.
Downsampling is the way to keep a long history cheaply: store coarse aggregates, say five-minute averages, for old data while raw resolution expires. Prometheus itself does not downsample, so this is a job for a long-term store behind remote write. The cheaper trick available in plain Prometheus is a recording rule, which precomputes an expensive aggregation once so dashboards read one series instead of thousands.
How do you design alerts that people trust?
The Prometheus alerting guidance summarises it in one sentence: keep alerting simple, alert on symptoms, have good consoles to pinpoint causes, and avoid pages where there is nothing to do. A symptom is what users feel, such as error ratio or latency. A cause is an internal state, such as CPU or a full connection pool. Page on the first, graph the second. The rule file below pairs a recording rule with a symptom alert that carries a for duration, and a down alert, because a dead target emits no series for a threshold to trip.
groups:
- name: api-symptoms
rules:
# Recording rule: compute the ratio once, reuse it in alerts and dashboards.
- record: job:http_errors:ratio_rate5m
expr: |
sum by (job) (rate(http_requests_total{job="api", status=~"5.."}[5m]))
/
sum by (job) (rate(http_requests_total{job="api"}[5m]))
# Symptom: users are getting errors. Not "CPU is high", which is a cause.
- alert: ApiHighErrorRatio
expr: job:http_errors:ratio_rate5m{job="api"} > 0.05
for: 10m # a 30-second blip never pages anyone
labels:
severity: page
annotations:
summary: "More than 5% of API requests failing for 10 minutes"
# Absence is also a symptom: a dead target produces no series, so a threshold never fires.
- alert: ApiTargetDown
expr: up{job="api"} == 0
for: 2m
labels:
severity: page
The for clause is what turns a flapping signal into a decision: the condition must hold continuously for ten minutes before the alert fires. The common mistake is the opposite pattern shown here.
# Wrong: a cause, with no for: clause. It pages on every compile spike and tells nobody what users feel.
- alert: HighCpu
expr: node_cpu_utilisation > 0.8
# Right: put the cause on a dashboard, page on the symptom above, and keep a "for:" on everything.
Before adding an alert, ask: if this fires at 3 a.m., what does the person do? If the answer is nothing, it belongs on a dashboard. Every alert that pages with no action teaches the team to ignore the next one, and that is how alert fatigue starts.
How do you scale a monitoring system past one server?
A single Prometheus server handles a surprising amount, and for one VPS it is the right answer. When it stops being enough, the options map to three different problems.
Federation: a higher-level Prometheus scrapes selected aggregates from lower-level ones, giving a global view without copying every series.
Remote write: each server streams its samples to a long-term, horizontally scalable store, which is also where downsampling and multi-year retention live.
Sharding by tenant or by service: separate servers per team or per customer, so one noisy tenant's cardinality cannot take down everyone's monitoring.
# Scaling out without changing the instrumentation: ship samples to a long-term store.
remote_write:
- url: https://metrics-store.internal/api/v1/push
queue_config:
max_samples_per_send: 2000 # batch size per request; tune against receiver limits
Remote write is the least invasive because the instrumentation, scrape config and alert rules stay as they are. Whatever you pick, keep the cardinality guardrails: scaling storage out raises the ceiling but does not change the multiplication.
A metrics system is a pipeline with a budget. Choose bounded labels, estimate series times samples times bytes before shipping a metric, keep raw data only as long as you will really query it, and page only on symptoms with a for duration. Do that and one Prometheus server stays boring for a long time, which is the point of monitoring.