Short answers to what readers ask most about this topic.
01How do you do back-of-the-envelope capacity estimation for a system?
State your assumptions first: daily active users, actions per user, record size, retention and replicas. Convert requests per day to QPS by dividing by 86,400, apply a peak multiplier, then derive storage, bandwidth, cache size and server count in turn. Treat the result as an order of magnitude and mark each guess so you can replace it with a measurement.
02How do I convert daily active users to QPS?
Multiply DAU by the requests each user makes per day, then divide by 86,400 seconds. For 2,000,000 DAU making 30 reads each, that is 60,000,000 reads a day, or 694.44 reads per second on average. Multiply by a peak multiplier, 3 in the worked example, to get about 2,083 per second at peak; the multiplier is an assumption you should replace with your own traffic data.
03How do I estimate storage for a system?
Multiply records per day by bytes per record, then by 365, the retention in years and the replication factor. In the worked example, 400,000 uploads a day at 1,151 KB each is 460.4 GB a day, 840 TB over five years and 2.52 PB with three copies. Check the largest field first, because it usually decides the storage technology.
04Is the 80/20 rule reliable for sizing a cache?
It is a heuristic borrowed from the Pareto principle, not a law, so use it as a starting point. Real access distributions vary, and a flatter one needs a larger cache for the same hit ratio. Size the cache for the hot 20 percent of daily reads, then log the real hit ratio after launch and replace the guess.
05Which latency numbers should I memorise, and are they still accurate?
Remember the ratios: memory is about 100 ns, a random SSD read about 150 us, a same-datacenter round trip about 500 us, a disk seek about 10 ms and a cross-ocean round trip about 150 ms. The popular list is dated about 2012, and hardware has changed since. Use it to reason about orders of magnitude, not as a current measurement.
Back-of-the-envelope capacity estimation turns a few stated assumptions into orders of magnitude. Daily active users give requests per day, dividing by 86,400 seconds gives average QPS, a peak multiplier gives peak QPS, and record size times retention times replicas gives storage. Bandwidth, cache size and server count follow the same way, each assumption flagged to measure later.
Someone says the service will have two million daily users and asks whether one server will do. The honest answer is a number, and the number takes about two minutes if you know the chain: users, requests, queries per second, bytes, cache, machines. Without the chain, the conversation turns into opinions about whether you need sharding.
This post is the generic method and a checklist you can reuse. It derives each step with the arithmetic shown, runs a TypeScript estimator on a photo-sharing example and prints its real output, and tabulates the powers of two and latency figures worth carrying in your head. The URL shortener, chat, news feed, web crawler and video streaming posts in this series each apply the method to their own inputs. Every number below is either derived from a stated assumption or cited in the sources at the end; none is a benchmark I measured.
What is back-of-the-envelope capacity estimation, and what is it for?
It is a way to get an answer within a factor of two or ten from assumptions you can state in one line each. Its job is to choose the class of architecture: whether one Postgres box and one Redis instance are enough, or whether you need sharding, a CDN and a fleet behind a load balancer. It is not forecasting and it is not a benchmark. A benchmark measures a running system; an estimate decides whether it is worth building the system that way.
The cheapest case is the one that tells you to stop. Take a carwash ERP issuing 200 receipts a day, assuming about 2 KB per receipt for illustration. 200 receipts / 86,400 seconds is 0.0023 writes per second, and 200 x 2 KB x 365 days is 146,000 KB, or 146 MB a year. Even at 100 times that volume, a single VPS holds it comfortably. The most useful output of an estimate is often the sentence that this does not need a distributed design, and it costs one minute.
Every input is a guess, and guesses multiply. Three inputs that are each off by 2x can leave the result off by 8x. Treat the output as an order of magnitude, and write the assumptions next to it so anyone can challenge the one that is wrong.
How do you convert daily active users into average and peak QPS?
Multiply daily active users (DAU) by the actions each user takes per day to get requests per day, then divide by the seconds in a day. A day is 86,400 seconds, which rounds to 10^5 for mental arithmetic with a 16 percent error, acceptable at this precision. The example service has 2,000,000 DAU, each viewing 30 photos and uploading 0.2 photos a day on average. That is 2,000,000 x 30 = 60,000,000 reads a day and 2,000,000 x 0.2 = 400,000 writes a day. Divided by 86,400, the averages are 694.44 reads per second and 4.63 writes per second.
The average does not size the fleet, because traffic is not flat. Multiply by a peak multiplier to get peak QPS. I use 3 here, and it is an assumption, not a constant: replace it with the ratio of your busiest hour to the daily average from your own metrics, and use a higher figure for sales, launches or paydays. With 3, peak reads are 694.44 x 3 = 2,083.33 per second and peak writes are 4.63 x 3 = 13.89 per second.
The read:write ratio is the number that tells you what to optimise. Here it is 60,000,000 / 400,000 = 150 to 1, which points at caches, read replicas and a CDN. A ratio near 1 or below points at the write path instead: append-friendly storage, batching and queues. Compute both rates separately, because the same DAU can describe two very different systems.
How do you estimate storage and bandwidth?
Storage is records per day x bytes per record x 365 x retention years x replication factor. For the photo service, assume 1,000 KB for the original, 150 KB for resized variants and 1 KB of metadata, so 1,151 KB per upload. 400,000 uploads x 1,151 KB = 460.4 GB a day. Over 5 years that is 460.4 GB x 365 x 5 = 840 TB, and with 3 copies (a common choice, and an assumption here) it is 2.52 PB. The metadata is 0.09 percent of the bytes, so the original file decides the storage choice: object storage, not rows in a database.
Bandwidth is peak QPS x bytes per request, in each direction. Ingress carries the upload and its variants: 13.89 uploads per second x 1,150 KB = 15.97 MB/s, which is 127.78 Mbps. Egress carries each viewed photo, assumed to be a 150 KB variant: 2,083.33 views per second x 150 KB = 312.5 MB/s. Network links are quoted in bits, so multiply bytes by 8: 312.5 MB/s is 2,500 Mbps, or 2.5 Gbps.
Compare that figure with the port you actually have. If your provider gives one server a 1 Gbps port (check the plan, it varies), 2.5 Gbps of egress means the answer is a CDN or object storage with its own edge, not a bigger application server. Bandwidth is where a design that looked fine on QPS alone most often breaks.
How big should the cache be, and is the 80/20 rule reliable?
A common starting point is to cache the hot 20 percent of what is read in a day, borrowing the 80/20 rule from the Pareto principle: roughly 80 percent of the requests land on 20 percent of the items. Treat it as a heuristic, not a law. Real access skew varies by workload, and a workload with a flatter distribution needs a bigger cache for the same hit ratio. For the photo service, metadata is 1 KB per read, so 0.2 x 60,000,000 reads x 1 KB = 12 GB, which fits in one Redis instance. The same method on the 150 KB image variants gives 0.2 x 60,000,000 x 150 KB = 1.8 TB, far too large for RAM, so images belong at a CDN edge or a disk cache.
This arithmetic multiplies reads, not distinct objects, so it overstates the cache when popular items repeat, which makes it a safe upper bound. If the cache then reaches an 80 percent hit ratio, only 20 percent of reads fall through: 2,083.33 x 0.2 = 416.67 reads per second reach the database at peak. That is the figure to size the database against, and it is five times lower than the raw peak.
After launch, log the real hit ratio and the real distinct-key count for a day. Replace the 20 percent guess with the measured number, and your cache estimate stops being a heuristic.
How many servers do you need?
Divide peak QPS by what one instance can serve, leaving headroom. The per-instance figure is the weakest number in the whole chain, because a handler that runs one indexed query and a handler that resizes an image differ by orders of magnitude. State it as an assumption and replace it with a load-test result as soon as you can run one. The example uses 500 requests per second per instance and a 50 percent target utilisation.
servers = ceil( peakQps / (perInstanceRps * targetUtilisation) )
// peakQps = 2,083.33 + 13.89 = 2,097.22 (from the QPS step)
// perInstanceRps = 500 ASSUMPTION - replace it with a load-test result
// targetUtilisation = 0.5 headroom for spikes, deploys and a dead node
servers = ceil( 2,097.22 / (500 * 0.5) ) = ceil( 8.39 ) = 9
// Cross-check with Little's law: in flight = arrival rate * time in system
// 2,097.22 req/s * 0.05 s (assumed 50 ms per request) = 104.86 concurrent requests
// 104.86 / 9 servers = about 12 in flight per server
The result is 9 application servers, or 10 if you want one spare so a node can die without leaving you under capacity. Little's law, the average number in a system equals the arrival rate times the time spent in it, gives a second view: about 105 requests in flight, roughly 12 per server. If that number is wildly different from your connection pool size or worker count, one of your assumptions is wrong, and the mismatch tells you which one to go and measure.
Which numbers should you memorise: powers of two, time and latency?
Storage arithmetic is faster if you know the powers of two cold. Each step up multiplies by 1,024, and for estimation you round it to a thousand and keep going. The names KB, MB and GB are used loosely for both 1,000 and 1,024 based units; the IEC names KiB, MiB and GiB are the precise binary ones.
Power
Exact bytes
Name and rounding
2^10
1,024
1 KiB, about a thousand bytes
2^20
1,048,576
1 MiB, about a million bytes
2^30
1,073,741,824
1 GiB, about a billion bytes
2^40
1,099,511,627,776
1 TiB, about a trillion bytes
2^50
1,125,899,906,842,624
1 PiB, about a quadrillion bytes
For time, one day is 86,400 seconds, close to 10^5. A year is 365 x 86,400 = 31,536,000 seconds, and one request per second is 86,400 a day or 2,592,000 in a 30-day month. For latency, the list usually attributed to Jeff Dean, originally by Peter Norvig, is labelled approximately 2012 in the gist that circulates. These are the lines that matter for estimation:
Main memory reference: 100 ns
Send 1K bytes over a 1 Gbps network: 10 us
Read 4K randomly from SSD: 150 us
Round trip within the same datacenter: 500 us
Read 1 MB sequentially from SSD: 1 ms
Disk seek: 10 ms
Send a packet from California to the Netherlands and back: 150 ms
Use them for ratios, not as current measurements. The gist dates them to about 2012, and Colin Scott's interactive page shows how the figures move by year; its code credits Norvig's 2002 numbers as the originals. Hardware has moved since, so quote the ratios: a memory reference at 100 ns against a random SSD read at 150 us is 1,500 times, and a disk seek at 10 ms against a datacenter round trip at 500 us is 20 times. Those ratios are why a cache hit changes a design, and they hold up better than any single absolute figure.
What does a full estimate look like end to end?
Here is the whole chain as a TypeScript estimator for the photo-sharing example. It has one function, one assumptions object and no dependencies, so each input is visible and the formulas match the arithmetic above.
The sixteen lines reduce to eight steps, each with its formula, so you can redo the arithmetic on paper and check the script against it:
Step
Formula
Photo-sharing result
Requests per day
DAU x actions per user
60,000,000 reads, 400,000 writes
Average QPS
requests per day / 86,400
694.44 reads, 4.63 writes
Peak QPS
average x peak multiplier of 3
2,083.33 reads, 13.89 writes
Storage per day
writes x 1,151 KB
460.40 GB
Storage over retention
per day x 365 x 5 years x 3 copies
840.23 TB raw, 2,520.69 TB replicated
Peak egress
peak read QPS x 150 KB x 8
312.50 MB/s, 2.50 Gbps
Metadata cache
0.2 x reads per day x 1 KB
12 GB
Application servers
peak QPS / (500 x 0.5)
9
Read the result as architecture, not as a spec. About 2,100 peak requests per second is modest. The 150 to 1 read ratio argues for a cache and a CDN. Roughly 2.5 petabytes argues for object storage. The 2.5 Gbps egress is the number that forces the CDN. Notice that none of these conclusions depends on the exact third digit; each would survive the inputs being wrong by 30 percent.
What is the checklist, and where do estimates go wrong?
Work the list in order and write each assumption down before you use it:
List the assumptions first: DAU, actions per user per day, record size, retention, replicas.
Convert to per-second figures with 86,400, and sanity-check against 10^5.
Apply a peak multiplier you can justify, and state its value.
Split reads from writes and note the ratio.
Compute storage with retention and replication, and check the units with the powers-of-two table.
Compute bandwidth in both directions, in bits, and compare it with the port or link you have.
Size the cache from a hot fraction, flag it as the 80/20 heuristic, and plan to measure the hit ratio.
Divide peak QPS by a per-instance throughput you mark as an assumption, with headroom for spikes and a dead node.
Four mistakes recur. Mixing bytes and bits turns 2.5 Gbps into 312 Gbps or 0.3 Gbps. Forgetting replication understates storage by the replica count. Sizing on the average instead of the peak leaves the service down at the exact moment it matters. And treating the estimate as a specification freezes guesses into a contract; revisit each flagged assumption once real traffic exists.
Write the assumptions, derive the chain, mark each guess, and keep the output as an order of magnitude. The point of the estimate is not the number but the decision it supports: one box or a fleet, a database or object storage, RAM or a CDN. When real metrics arrive, replace the flagged assumptions one by one, starting with the peak multiplier and the per-instance throughput.