Short answers to what readers ask most about this topic.
01How do you approach a system design interview step by step?
Work through seven steps in order: clarify requirements and scope, estimate scale, define the API and data model, draw the high-level design, deep-dive the riskiest component, discuss bottlenecks, failure modes and trade-offs, then summarise. Each step produces something the next one uses, so skipping one makes the rest vague. Say your assumptions aloud and attach a number to each decision.
02What questions should I ask the interviewer in a system design interview?
Ask about the core use cases, the read-to-write ratio, the expected number of users and their growth, the latency and availability expected, whether stale data is acceptable, and how long data is kept. Also ask what is out of scope so you do not design accounts, search and editing by accident. If the interviewer will not answer, state an assumption and continue.
03How long should each part of a 45-minute system design interview take?
There is no official split, but a sensible starting point is about 5 minutes for requirements, 5 for estimation, 7 for API and data model, 7 for the high-level design, 10 for the deep dive, 7 for bottlenecks and trade-offs, and 4 to wrap up. Treat it as a suggestion, since the interviewer may steer you toward one area. If you run long, shorten the deep dive rather than skipping failure modes.
04What are the most common mistakes in a system design interview?
The most common are naming technologies before the requirements are clear, giving no numbers for scale, and drawing a design in which nothing ever fails. Others are designing for the wrong scale, such as sharding data that fits on one disk, and talking for ten minutes without checking the interviewer's intent. Each has a cheap fix: describe the need first, run the arithmetic, and walk the request path asking what happens when each hop is slow or down.
05Do I need to do back-of-the-envelope estimation in every system design interview?
Yes, in some form, because the numbers decide which components you need. Even a rough result such as about a dozen writes and about a hundred reads per second tells you a single database node is enough and that storage growth is the real concern. Round aggressively, state every assumption, and stop once you have requests per second, storage per year and bandwidth.
System Design Interview Framework: A Step-by-Step Approach
How to approach a system design interview step by step: requirements, scale estimates, API and data model, diagram, deep dive, trade-offs, with a worked pastebin example.
To approach a system design interview step by step, clarify functional and non-functional requirements, estimate scale with simple arithmetic, define the API and data model, sketch a high-level diagram, deep-dive the riskiest component, then discuss bottlenecks, failure modes and trade-offs before summarising. State assumptions aloud and attach a number to every decision.
Most people who stall in a system design interview do not lack knowledge. They lack an order of operations. A prompt such as design a pastebin is open-ended on purpose, and an open prompt with no procedure turns into a list of technology names recited at the whiteboard.
I should be plain about the authority here. I am writing this from the candidate's side of the table, not as someone who runs interviews, so there are no interview anecdotes below. It is a checklist assembled from the standard literature: the definition of systems design, the Google SRE book, the AWS Well-Architected Framework and the Azure Architecture Guide. The worked example at the end uses only numbers I derive in front of you.
What is a repeatable framework for a system design interview?
A framework is a fixed order of seven steps. Each step produces something the next step consumes, which is why skipping one makes the later ones vague.
Clarify requirements and scope: separate what the system does (functional) from how well it must do it (non-functional), and agree what is out of scope.
Estimate scale: turn users and actions into requests per second, storage and bandwidth, using rounded arithmetic.
Define the API and data model: the handful of operations and the entities they touch, before any box is drawn.
Draw the high-level design: clients, services, stores and the path a request takes through them.
Deep-dive the riskiest component: the part where the numbers or the requirements make the design hard, not the part you find most fun.
Examine bottlenecks, failure modes and trade-offs: what breaks first, what happens when it does, and what each choice cost.
Wrap up: summarise the design against the requirements and name what you would build next.
The definition of systems design on Wikipedia is the process of defining the architecture, modules, interfaces and data of a system to satisfy specified requirements. That sentence is the framework in miniature: requirements first, then interfaces and data, then architecture. Interviewers are watching whether you work in that direction, not whether you land on a particular diagram.
How do you clarify requirements and scope?
Spend the first minutes asking, and say your assumptions out loud when nobody answers. The aim is to split the prompt into functional requirements, which are features, and non-functional requirements, which are the qualities that decide the architecture. The table shows the questions that do the most work.
Kind
Question to ask
What the answer changes
Functional
Who uses it and what are the two or three core actions?
The API surface and the entities in the data model
Functional
What is explicitly out of scope: accounts, search, editing?
How many components you draw, and how much time the deep dive gets
Non-functional
What latency and availability does the user expect?
Caching, replication and whether a single node is acceptable
How many users, how much data, how long is it kept?
Every number in the estimation step
The Google SRE book gives the vocabulary for the non-functional half: it names latency, availability and throughput among the indicators a service is judged by, and it recommends stating targets as percentiles rather than averages. Borrow that. Saying the read path should answer in under 200 ms at the 99th percentile is a requirement; saying it should be fast is not.
Questions worth asking in almost any prompt:
What are the core use cases, and which one matters most?
Is the workload read-heavy or write-heavy, and by what ratio?
What is the expected number of users, and how fast might it grow?
Does the system need to be available during a failure, or is a short outage tolerable?
Are there data retention, deletion or privacy constraints?
Which parts may I treat as out of scope or as an existing service?
Write the agreed requirements in one corner of the board and leave them there. When you reach the trade-off step, you point at a requirement instead of defending a preference.
How do you estimate scale without a calculator?
Estimation exists to choose an architecture, not to size a bill. The capacity-estimation post in this series covers the technique in depth; the short version is to round hard, keep every assumption visible, and finish with four numbers: writes per second, reads per second, storage per year and bandwidth. If the answer is a dozen requests per second, you have just learned that one database node is not the hard part.
Pick round constants once and reuse them. A day has 86,400 seconds, which you round to about 100,000 when speed matters. Peak traffic is a multiple of the average, and two to three times is a reasonable stated assumption as long as you say it is an assumption. Two minutes of this arithmetic prevents ten minutes of designing for the wrong scale.
How should you spend a 45-minute slot?
This is a suggestion, not a rule. Interviews differ in length and format, and an interviewer who wants to dig into one area will override any schedule you carry in. The split below is a starting point I find sensible given that the deep dive and the trade-offs are where the differences between candidates show up.
Step
Suggested minutes
What you should have on the board afterwards
1. Clarify requirements
5
Functional list, non-functional targets, out of scope
2. Estimate scale
5
Requests per second, storage per year, bandwidth
3. API and data model
7
Endpoints and one or two tables or entities
4. High-level design
7
A box-and-arrow diagram with the request path
5. Deep dive
10
One component worked through in detail
6. Bottlenecks and trade-offs
7
Failure modes and the cost of each choice
7. Wrap up
4
A one-minute summary and next steps
The total is 45. If you are running long at the high-level design, shorten the deep dive instead of skipping the failure discussion; a design nobody has stress-tested is worth less than a design with one component explored thoroughly.
How do you discuss trade-offs and failure modes?
A trade-off statement has three parts: the choice, what it buys, and what it costs. Tie it to a requirement or a number you already wrote down. These phrases keep the discussion structured and stop it from sounding like a preference.
I would choose this because the requirement we agreed is a read latency target, and this option serves reads from memory.
This buys us availability at the cost of occasionally stale reads, and we said stale reads were acceptable for a few seconds.
If this component fails, users see this specific symptom, and we recover by doing this specific thing.
The estimate says we do not need this yet, so I would leave it out and revisit it at ten times the load.
This is the point where I would measure before deciding, and here is the metric I would look at.
For failure modes, walk the request path from the diagram and ask what happens when each hop is slow, down or wrong. The Google SRE chapter on cascading failures explains why slow is the dangerous case: an overloaded server that answers slowly invites retries, and retries add load to the server that is already struggling. The CAP-theorem and cache-aside posts in this series are good rehearsal material for the consistency and staleness side of the same conversation.
Do not offer a retry as your only answer to a failing dependency. Unbounded retries are a documented way to turn a slow dependency into a full outage, so pair any retry with a limit, backoff and a way to shed load.
What are the most common system design interview mistakes?
Three mistakes account for most of the weak answers, and all three are about order rather than knowledge: naming tools too early, skipping numbers and ignoring failure. The table pairs each with how it shows up and the correction, and adds two related ones.
Mistake
How it shows up
Fix
Jumping to technology names
Kafka or Redis appears in the first minute, before anyone knows the read rate
Describe the need first, such as a buffer or a fast read path, then name a tool and the alternative you rejected
No numbers
Words like huge scale, millions of users, low latency with nothing attached
Run the estimation step and let the result veto or justify each component
Ignoring failure
A diagram in which every box always works
Walk the request path and say what each hop does when it is slow or down
Designing for the wrong scale
Sharding a dataset that fits on one disk
Compare the estimate to what a single node handles before adding machines
Silent monologue
Ten minutes without checking the requirements are the ones the interviewer meant
Pause after each step and ask whether to go deeper or move on
The fix for the first mistake is the cheapest to practise: before you say a product name, finish the sentence the system needs a thing that does X at Y per second. If you cannot fill in Y, you have skipped step two.
What does the framework look like on a real prompt: design a pastebin?
Here is the whole framework applied end to end to a pastebin, a service where a user submits text and gets a short link back. Every figure below is an assumption I state, and every other number is derived from those assumptions. None is a measurement.
Requirement
Decision or assumption
Functional: create
Submit text, receive a short link
Functional: read
Open the link, see the text; an optional expiry time
Out of scope
Accounts, editing, search, syntax-aware diffing
Non-functional: scale
Assume 1 million new pastes a day, 10 reads per write, 10 KB average, 512 KB maximum
Non-functional: latency
Read under 200 ms at the 99th percentile
Non-functional: availability and retention
Reads stay up through a single node failure; pastes are kept for one year
Step two is the arithmetic. It shows that traffic is tiny and storage is the real growth axis, which is what tells us where to spend the deep dive.
// Assumptions (stated aloud): 1M new pastes/day, 10 reads per write,
// 10 KB average body, 200 B metadata row, 1 year retention, peak = 3x average.
const SECONDS_PER_DAY = 86_400;
const writesPerSec = 1_000_000 / SECONDS_PER_DAY; // 11.6 -> call it 12
const readsPerSec = 10_000_000 / SECONDS_PER_DAY; // 115.7 -> call it 116
const peakWrites = 3 * writesPerSec; // ~35
const peakReads = 3 * readsPerSec; // ~350
const bodyBytesPerYear = 365_000_000 * 10_000; // 3.65e12 = 3.65 TB
const metaBytesPerYear = 365_000_000 * 200; // 7.3e10 = 73 GB
const readBandwidth = readsPerSec * 10_000; // ~1.16 MB/s average
// Key space: 7 base62 characters.
const keySpace = 62 ** 7; // 3_521_614_606_208 ~ 3.5e12
const fillAfter10y = 3_650_000_000 / keySpace; // ~0.001 -> 0.1% full
// Reading: ~350 reads/s at peak is one modest database node.
// The real growth axis is storage, so that is where the deep dive goes.
Step three: the API sketch and the data model. The body of a paste lives in an object store, and the database holds only a small metadata row pointing to it, because the metadata is about 200 bytes and the body is about 10 KB.
POST /v1/pastes
body: { "content": "...", "syntax": "plain", "ttl_seconds": 86400 } // ttl optional
201: { "id": "aZ3k9Qp", "url": "https://paste.example/aZ3k9Qp", "expires_at": "2026-10-11T08:00:00Z" }
413: content over 512 KB
429: create rate limit exceeded
GET /v1/pastes/aZ3k9Qp
200: { "id": "aZ3k9Qp", "content": "...", "syntax": "plain", "created_at": "..." }
404: unknown id
410: paste existed and has expired
-- Metadata only. The body lives in the object store under blob_key.
CREATE TABLE pastes (
id char(7) PRIMARY KEY, -- base62; 62^7 ~ 3.5e12 keys
blob_key text NOT NULL, -- object-store key for the body
size_bytes integer NOT NULL CHECK (size_bytes BETWEEN 1 AND 524288),
syntax text NOT NULL DEFAULT 'plain',
created_at timestamptz NOT NULL DEFAULT now(),
expires_at timestamptz -- NULL = keep for the retention period
);
-- The expiry sweeper scans this; partial so non-expiring rows cost nothing.
CREATE INDEX pastes_expires_idx ON pastes (expires_at)
WHERE expires_at IS NOT NULL;
-- Random key + unique primary key: a collision fails the insert, the app retries with a new key.
-- Write this row LAST, after the body is stored, so a failure leaves an orphan blob, never a dangling row.
Step four, the high-level design, is short for this prompt: a client talks to a stateless web service behind a load balancer; the service writes metadata to a relational database and the paste body to an object store; reads check a cache first. Step five picks the riskiest part, which is not throughput, since a single database handles about a hundred reads per second easily. It is key generation and unbounded storage growth.
For keys, seven base62 characters give about 3.5 trillion values. Picking a random key and relying on a unique constraint means a collision on insert is retried with a new key. After ten years of pastes the table would be about 0.1 percent full, so the expected number of retries stays negligible. A counter-based scheme avoids retries but needs a coordinated counter and makes keys guessable, so name that trade-off aloud and say which one you chose and why.
Steps six and seven. Bottlenecks: storage grows by about 3.65 TB a year, so the object store, not the database, is the thing to watch, and expiry must actually delete bodies. Failure modes: if the cache is down, reads fall through to the database, which can absorb the load at this scale, so the cache is an optimisation and not a dependency. If the object store write succeeds and the database insert fails, you have an orphaned blob, so write the metadata row last and sweep for orphans. Wrap up by restating the requirements and saying what is next: abuse prevention and rate limiting on the create endpoint.
How do you practise and wrap up the interview?
The last four minutes are a short summary, not new material. Restate the requirements in one sentence, name the component you explored and the main trade-off, and list one or two things you would add with more time. As a final pass, the AWS Well-Architected Framework offers a checklist of six pillars: operational excellence, security, reliability, performance efficiency, cost optimization and sustainability. Asking which of the six you have not touched is a quick way to find a gap.
To practise, run the seven steps on one prompt per sitting, set a timer, and write down the numbers before the diagram. The other posts in this series are drills with the full reasoning shown: designing a URL shortener, the rate limiting algorithms post, cache-aside versus write-through caching, and the CAP theorem with practical examples. For a gentler start, the system design explained for developers post covers the basics, and the Azure Architecture Guide is a catalogue of architecture styles and design patterns worth skimming for vocabulary.
The rule I take from all of this is that order beats cleverness. Ask before you draw, count before you choose, and say what each choice costs. A plain design with numbers behind it and its failure modes named will usually read better than an elaborate one that never mentions how it breaks.