Short answers to what readers ask most about this topic.
01What is system design in software development?
System design is deciding how a system's components, data and interfaces fit together so it meets requirements for load, reliability and cost. The unit of work is a decision with a trade-off, such as a cache that speeds reads but can serve stale data. It happens before and alongside coding, not instead of it.
02How do I start learning system design from scratch?
Start with the vocabulary: latency, throughput, availability, consistency and idempotency. Then read one in-depth book such as Designing Data-Intensive Applications, and apply the ideas to a system you already know. Practise by writing a one-page design note for a small feature before you code it.
03Do I need system design skills as a junior developer?
You do not need to design large distributed systems, but you benefit from understanding why the system you work on is built as it is. Estimating load and asking what happens on failure are useful from your first projects. The depth grows naturally as you own bigger features.
04How do you estimate load in a system design?
Multiply users or outlets by actions per day, divide by 86,400 seconds for the average rate, then apply a stated peak factor. Add row size times volume for storage, and use Little's law to find requests in flight. Write every assumption down so others can challenge it.
05When should I add a cache, a queue or a load balancer?
Add one only when you can measure the symptom it solves: repeated reads for a cache, slow work blocking requests for a queue, a saturated CPU for more instances behind a load balancer. Each component adds a cost such as stale data, eventual results or stateless servers. If no symptom exists yet, adding nothing is a valid design.
System design is the process of deciding how a software system's components, data and interfaces fit together to meet stated requirements for load, reliability and cost. You learn it by estimating load with arithmetic, naming the failure modes, choosing the simplest building blocks that survive them, and writing the trade-offs down.
The first time someone asked me to draw the architecture before writing any code, I drew five boxes and three arrows and felt finished. Then they asked what happens when the database host dies at peak hour, and the drawing had no answer.
This post answers the question people search for: what system design is and how to learn it. It leans on cited definitions for the vocabulary and on arithmetic you can check for everything else. The example is a hypothetical multi-outlet POS, the kind of problem I work near with ERP, POS and NestJS on Postgres, Redis and Docker. Every input number is labelled as an assumption, not a measurement.
What is system design in software development?
System design is the work of defining the parts of a system, how they talk to each other, and what data they hold, so the whole thing meets requirements you can state. The Wikipedia entry on systems design frames it as understanding component parts and how they interact. In software that means services, databases, queues, caches and the contracts between them.
It is different from writing code in one respect: the unit of work is a decision, not a function. Each decision has a cost somewhere else. A cache makes reads fast and makes data stale. A queue makes requests quick and makes results eventual. System design is choosing which cost you can live with, before you have paid it in production.
What do you actually decide when you design a system?
Five decisions cover most of it, and they come in this order because each one constrains the next.
Requirements: what must it do, and how fast, how available and how correct must it be. Unstated requirements are where designs quietly fail.
Load: how many requests per second, how much data per year, and how bursty. This is arithmetic, not opinion.
Data model: what the entities are, which store owns each one, and what must be consistent at the same instant.
Interfaces: the API shape, including what a client may safely retry.
Failure: what breaks first, what the user sees, and how much loss is acceptable. Cloud guidance such as the AWS Well-Architected Framework treats reliability, performance efficiency, cost optimisation and security as separate pillars precisely because they pull against each other.
Notice that scaling is not on the list as its own item. It falls out of the load estimate. Many designs fail by answering the scaling question before measuring whether there is one.
How do you estimate load before drawing any boxes?
Write the assumptions as code or a table and let the arithmetic speak. Below is a hypothetical rollout of 200 outlets at 300 transactions each per day. The peak factor, row size and service time are assumptions I chose for illustration; replace them with your own numbers as soon as you have real ones.
// Back-of-envelope sizing for a hypothetical multi-outlet POS.
// Every input is an ASSUMPTION you state out loud; every output is arithmetic.
const outlets = 200;
const txPerOutletPerDay = 300;
const secondsPerDay = 86_400;
const peakFactor = 10; // assumption: the busiest second is ~10x the mean
const bytesPerTx = 2_000; // assumption: header + ~6 lines + payments, as rows
const serviceTimeSeconds = 0.05; // assumption: 50 ms of work per request
const txPerDay = outlets * txPerOutletPerDay; // 60,000
const meanTps = txPerDay / secondsPerDay; // 0.694...
const peakTps = meanTps * peakFactor; // 6.94...
// Little's law: L = lambda * W (items in the system = arrival rate * time each spends in it)
const inFlight = peakTps * serviceTimeSeconds; // 0.347... requests in flight
const gbPerYear = (txPerDay * bytesPerTx * 365) / 1_000_000_000; // 43.8 GB
// Reading: ~7 writes/s at peak and ~44 GB/year is a single-Postgres problem.
// The design question is NOT "how do we scale?" but "what happens when it fails?"
The result is the useful part. About 7 writes per second at peak, under half a request in flight by Little's law (in-flight work equals arrival rate times time spent in the system), and roughly 44 GB a year. That fits one well-configured Postgres instance. So the interesting design problems are backups, retries and failover, not sharding. Had the arithmetic said 7,000 writes per second, the design would look completely different, and that is the point of doing it first.
Round aggressively and state the assumption next to the number. A reviewer can argue with a 10x peak factor in seconds. They cannot argue with a vague claim that the system must scale.
Which building block solves which problem?
Components are answers to measured symptoms, not decorations. Use this table as a decision checklist: find the symptom you can actually observe, then look at what the fix costs you.
Symptom you can measure
Building block
What it costs you
The same rows are read again and again and the database is the bottleneck
Cache such as Redis in front of the database
Stale data and invalidation logic
Slow work (PDF, email, third-party sync) blocks the HTTP request
Queue plus a worker process
Eventual results, retries, duplicate handling
One application process saturates the CPU
More instances behind a load balancer
Servers must be stateless, sessions move out of memory
A query gets slower as a table grows
An index first, a read replica later
Slower writes, and replica lag on reads
One database host is a single point of failure
Replication with failover, plus tested backups
Operational complexity and possible data loss window on failover
Read the right column twice. Every row buys one property by spending another, and a design that lists only benefits is a sales pitch, not a design. If none of the left-hand symptoms is true yet, the correct number of new components is zero. On a single VPS running NestJS, Postgres and Redis in Docker, that is often the honest answer.
How do I learn system design as a developer?
Learn it as a loop of estimate, design, break and revise, not as a catalogue of famous architectures. A practical order:
Learn the vocabulary from a neutral source: latency, throughput, availability, consistency, idempotency. Wikipedia and the cloud vendors' architecture guides are good enough to start.
Read one deep book end to end. Designing Data-Intensive Applications by Martin Kleppmann covers storage, replication, partitioning and consistency in a way that explains why the building blocks behave as they do.
Take a system you already work on and run the load arithmetic on it. Real numbers from your own logs beat any textbook example.
Write a one-page design note for a small feature before you code it, including a rejected alternative and a failure mode.
Break your own design on paper: kill each component in turn and write down what the user sees.
Interview-style practice has a place, but it trains speed at drawing boxes. The note-and-break loop trains the judgement that real work asks for.
What does a one-page design note look like?
Short enough to read in five minutes, specific enough to be wrong. This is the template I would use for the POS sync feature sized above. The RPO and RTO figures are targets a team agrees on, not measurements.
# Design note: receipt sync for outlet POS (one page, written BEFORE code)
## 1. Requirements
Functional : an outlet can sell while offline; sales reach head office within 5 min of reconnect.
Non-func : peak ~7 writes/s (see sizing); no lost sale; no double-posted sale.
## 2. Data
sale(id uuid PK, outlet_id, total_idr bigint, created_at timestamptz) -- money as integer rupiah
sale_line(sale_id FK, sku, qty, price_idr)
## 3. Interface
POST /v1/sales Idempotency-Key: <sale uuid> -> 201 first time, 200 on replay
## 4. Failure modes
Network drops mid-request -> client retries -> idempotency key makes the retry safe.
DB host dies -> restore from backup; accepted RPO 15 min, RTO 1 h.
## 5. Rejected alternatives
Kafka : one consumer, ~7 msg/s; operating it costs more than the problem.
Microservices: one team; a network hop buys nothing yet.
## 6. Revisit when
peak > 200 writes/s, or a second team owns part of the schema.
Two sections do most of the work. Failure modes forces you to say what happens on a retry, which is why the interface line carries an idempotency key. Rejected alternatives records that you considered Kafka or microservices and said no for a reason tied to the load numbers. Six months later that section saves an argument.
What mistakes do beginners make in system design?
The most common one is designing for a load that does not exist. Reaching for sharding, Kafka or a service mesh on the strength of a famous company's architecture, with no estimate of your own, adds operating cost that the sizing above shows you do not need.
Do not copy a large company's architecture because it appears in a talk. It was shaped by their load, team size and failure history. Your arithmetic should justify every component, and a component with no measured symptom behind it is a liability.
The second mistake is ignoring failure and asking only how to make it fast. The third is leaving trade-offs unwritten, so the next engineer cannot tell a deliberate choice from an accident. A design that states its limits, and the conditions under which you would revisit them, is easier to trust than one that claims to scale to everything.
System design is a habit of deciding with numbers and writing down what each decision costs. Estimate the load, name the failure, add a component only when a measured symptom demands it, and record the alternative you rejected. Do that on a real system you already know and you are practising the skill itself.