Short answers to what readers ask most about this topic.
01How does video streaming work?
The video is transcoded once into several resolutions and bitrates, and each version is cut into segments of a few seconds. The segments are stored as ordinary files and served over HTTP, usually through a CDN. A playlist lists them, and the player downloads one segment at a time while choosing the quality it can sustain.
02What is HLS and what is an .m3u8 file?
HLS is HTTP Live Streaming, defined in RFC 8216. An .m3u8 file is a UTF-8 text playlist: a master playlist lists the available renditions with their BANDWIDTH and RESOLUTION, and each media playlist lists the segment files in order with their durations.
03How does adaptive bitrate streaming choose quality?
The player measures how fast each segment downloads, smooths that into a throughput estimate, and picks the highest rendition whose declared bandwidth fits under it with a safety margin. It also watches the buffer and drops a rung when the buffer runs low. Switching happens only at segment boundaries.
04How much storage and bandwidth does video streaming need?
It depends on the ladder, but it is simple arithmetic: bitrate times duration divided by 8 gives bytes. With a four-rung ladder of 5,128, 2,928, 1,528 and 928 kbps, one hour of video takes 4.73 GB, and 10,000 viewers at 2,928 kbps need 29.28 Gbps of egress. Egress, not storage, is what forces a CDN.
05What is the difference between live streaming and VOD?
VOD uses a finished playlist that ends with EXT-X-ENDLIST and never changes, so it caches for a long time. Live keeps appending segments and the player reloads the playlist, so its cache lifetime must be about one segment. Live also encodes in real time and runs several segment durations behind the camera.
Design a Video Streaming System: HLS, CDN and Adaptive Bitrate
How video streaming works and how to design it: a transcode ladder, HLS segments and manifests, object storage, a CDN, adaptive bitrate logic, signed URLs and live versus VOD, with the storage and egress arithmetic shown.
Video streaming works by transcoding every upload into several resolutions and bitrates, cutting each into segments of a few seconds, storing them as plain files in object storage, and serving them over HTTP through a CDN. A manifest lists the renditions, and the player measures its download speed and switches quality between segments.
The question that makes video different from every other system design exercise is that nothing is streamed in the way the word suggests. There is no long-lived connection pushing frames. There is a folder of small files, a text playlist, and a player that keeps asking for the next file.
This post designs that pipeline end to end: upload, transcode ladder, segmenting, storage, CDN, the player's quality logic, access control and the live variant. The numbers are worked from stated assumptions, not measured, and the playlist and ffmpeg command are checked against RFC 8216 and the ffmpeg documentation. I run ERP and POS workloads on a single VPS, so I also say where that stops working.
How does video streaming actually work?
A streaming platform is a file pipeline with a smart client at the end. The video is converted once, ahead of time, into many small files at several quality levels, and ordinary HTTP servers hand those files out. The path looks like this:
Upload: the original file lands in object storage, usually through a presigned upload so it never passes through your API servers.
Transcode: a worker encodes the original into a ladder of renditions, each a different resolution and bitrate.
Segment: each rendition is cut into segments of a few seconds, and a playlist lists them in order.
Store and distribute: segments and playlists are written to object storage and fronted by a CDN, which caches them close to viewers.
Play: the player downloads the master playlist, picks a rendition, then fetches segments one at a time, re-choosing the rendition as conditions change.
Because the delivery is plain HTTP, the CDN needs no video-specific logic. A segment is just a cacheable file, which is the whole reason this design scales. Three delivery formats are worth telling apart:
Format
How it adapts
Trade-off
Progressive MP4
It does not. One file, one quality, fetched with range requests.
Simplest possible, fine for short clips, but a slow connection stalls or you pick one quality for everyone.
HLS
A .m3u8 playlist lists renditions and segments; the player switches between them. Defined in RFC 8216.
Plays natively on Apple platforms and through JavaScript players elsewhere. Segments are MPEG-TS or fragmented MP4.
MPEG-DASH
An XML .mpd manifest describes the renditions; the player switches between them.
An open standard with the same segment-and-manifest idea, played in browsers through Media Source Extensions.
For a small product I would pick HLS: one format, one set of files, and the player libraries are mature. Packaging the same segments for DASH as well is a later optimisation, not a day-one need.
How much storage and bandwidth does a video platform need?
Do this arithmetic before choosing any infrastructure, because egress, not storage, decides the design. The inputs below are assumptions for a design exercise, and each line is derived from the one above it.
Assumed inputs (a design exercise, not a measurement):
ladder, video + 128 kbps AAC audio:
1080p 5,000 + 128 = 5,128 kbps
720p 2,800 + 128 = 2,928 kbps
480p 1,400 + 128 = 1,528 kbps
360p 800 + 128 = 928 kbps
segment length = 6 s
catalogue = 1,000 hours of source video
concurrent viewers at peak = 10,000, all on the 720p rung
origin network port = 1 Gbps (assumed single VPS)
Storage (kbps x 3,600 s / 8 = kB per hour)
1080p 5,128 x 3,600 / 8 = 2,307,600 kB = 2.31 GB per hour
720p 2,928 x 3,600 / 8 = 1,317,600 kB = 1.32 GB per hour
480p 1,528 x 3,600 / 8 = 687,600 kB = 0.69 GB per hour
360p 928 x 3,600 / 8 = 417,600 kB = 0.42 GB per hour
whole ladder: 10,512 x 3,600 / 8 = 4,730,400 kB = 4.73 GB per hour of video
catalogue: 4.73 GB x 1,000 = 4.73 TB
Objects
segments per rendition per hour = 3,600 / 6 = 600
per hour of video = 600 x 4 renditions = 2,400 files
catalogue = 2,400 x 1,000 = 2,400,000 files
Egress at peak
10,000 x 2,928 kbps = 29,280,000 kbps = 29.28 Gbps
per hour of viewing = 10,000 x 1.3176 GB = 13,176 GB = about 13.2 TB
machines to push that at line rate from a 1 Gbps port = 29.28 / 1 -> 30
Origin load behind a CDN (origin sees only the misses)
95% hit ratio -> 5% x 29.28 Gbps = 1.464 Gbps (still over one port)
99% hit ratio -> 1% x 29.28 Gbps = 0.293 Gbps (fits)
Request rate for 10,000 viewers
6 s segments : 10,000 / 6 = 1,667 segment requests per second
2 s segments : 10,000 / 2 = 5,000 segment requests per second
Two results matter. Storage is cheap by comparison: the whole 1,000-hour catalogue fits in 4.73 TB, and the 2.4 million files are an object-count problem, not a capacity problem. Egress is the real constraint: 10,000 concurrent viewers need 29.28 Gbps, which is about thirty one-gigabit ports of sustained output.
This is why a CDN is not optional here. Even at a 95 percent hit ratio the origin still sees 1.464 Gbps, more than one port, so the hit ratio you need is closer to 99 percent. Segments help: a popular video's segment is requested by thousands of viewers, so it is fetched from origin once per edge location and served from cache afterwards. For how edge caching behaves, see the post on how a CDN works.
Do not serve segments from your application server or a single VPS and call it streaming. The arithmetic above says 10,000 viewers at 720p need 29.28 Gbps; one 1 Gbps port tops out at about 341 such viewers (1,000,000 kbps divided by 2,928 kbps). Put object storage behind a CDN from day one.
How do you transcode into a bitrate ladder?
A ladder is the set of renditions the player can choose between. The rungs below are my assumed starting point, not a standard: 1080p at 5,000 kbps, 720p at 2,800, 480p at 1,400 and 360p at 800, each with 128 kbps AAC audio. Tune them to your content, since screen recordings and sport compress very differently. One ffmpeg command can produce the whole ladder and the playlists.
# One input, four renditions, HLS output. Assumes a 24 fps source.
# -g 144 = 6 s x 24 fps, so every segment starts on a keyframe.
# -sc_threshold 0 stops x264 inserting extra keyframes at scene cuts,
# which would shift segment boundaries differently per rendition.
ffmpeg -i input.mp4 \
-filter_complex "[0:v]split=4[a][b][c][d];[a]scale=-2:1080[v1080];[b]scale=-2:720[v720];[c]scale=-2:480[v480];[d]scale=-2:360[v360]" \
-map "[v1080]" -map "[v720]" -map "[v480]" -map "[v360]" \
-map 0:a:0 -map 0:a:0 -map 0:a:0 -map 0:a:0 \
-c:v libx264 -preset medium -g 144 -keyint_min 144 -sc_threshold 0 \
-b:v:0 5000k -maxrate:v:0 5350k -bufsize:v:0 7500k \
-b:v:1 2800k -maxrate:v:1 2996k -bufsize:v:1 4200k \
-b:v:2 1400k -maxrate:v:2 1498k -bufsize:v:2 2100k \
-b:v:3 800k -maxrate:v:3 856k -bufsize:v:3 1200k \
-c:a aac -b:a 128k \
-f hls -hls_time 6 -hls_playlist_type vod \
-hls_flags independent_segments -hls_segment_type mpegts \
-hls_segment_filename "out/%v/seg_%03d.ts" \
-master_pl_name master.m3u8 \
-var_stream_map "v:0,a:0,name:1080p v:1,a:1,name:720p v:2,a:2,name:480p v:3,a:3,name:360p" \
"out/%v/index.m3u8"
# Result: out/master.m3u8, out/1080p/index.m3u8, out/1080p/seg_000.ts ... and so on per rung.
Three details carry the design. First, -g 144 with -sc_threshold 0 forces a keyframe every 6 seconds at 24 fps, and the HLS muxer cuts a segment only on the next keyframe after hls_time has passed, so the keyframe interval must match the segment length or segments will not be the length you asked for. Second, every rendition must cut at the same instants, which is what lets the player switch between them at a segment boundary without a visible jump. Third, -maxrate and -bufsize cap the bitrate spikes that the playlist's BANDWIDTH attribute has to promise.
Transcoding is the expensive, bursty step, so treat it as a queue of jobs, not a request handler. Put the original in object storage, enqueue a job, run the encoder in a worker container, and write the output back. A failed job retries from the original and never touches what viewers are already watching.
Keep the original forever and treat every rendition as a cache of it. When you later add an AV1 or HEVC ladder, or change the segment length, you re-run the job from the source instead of re-encoding an already-compressed rendition and losing quality twice.
What does an HLS manifest (.m3u8) look like?
There are two kinds of playlist. The master playlist lists the renditions, one EXT-X-STREAM-INF line each, and RFC 8216 requires the BANDWIDTH attribute, the peak segment bit rate in bits per second. AVERAGE-BANDWIDTH, RESOLUTION, FRAME-RATE and CODECS are optional but worth setting so the player can rule out renditions it cannot decode.
#EXTM3U
#EXT-X-VERSION:3
#EXT-X-INDEPENDENT-SEGMENTS
#EXT-X-STREAM-INF:BANDWIDTH=984000,AVERAGE-BANDWIDTH=928000,RESOLUTION=640x360,FRAME-RATE=24.000,CODECS="avc1.42c01e,mp4a.40.2"
360p/index.m3u8
#EXT-X-STREAM-INF:BANDWIDTH=1626000,AVERAGE-BANDWIDTH=1528000,RESOLUTION=854x480,FRAME-RATE=24.000,CODECS="avc1.4d401e,mp4a.40.2"
480p/index.m3u8
#EXT-X-STREAM-INF:BANDWIDTH=3124000,AVERAGE-BANDWIDTH=2928000,RESOLUTION=1280x720,FRAME-RATE=24.000,CODECS="avc1.64001f,mp4a.40.2"
720p/index.m3u8
#EXT-X-STREAM-INF:BANDWIDTH=5478000,AVERAGE-BANDWIDTH=5128000,RESOLUTION=1920x1080,FRAME-RATE=24.000,CODECS="avc1.640028,mp4a.40.2"
1080p/index.m3u8
# BANDWIDTH = maxrate + 128k audio (peak). AVERAGE-BANDWIDTH = the target bitrate + audio.
# CODECS must match what the encoder really wrote: confirm with ffprobe, do not copy these blindly.
Each rendition has its own media playlist: a target duration, then an EXTINF duration and URI for every segment. The final segment is shorter because the video does not divide into 6 seconds exactly. The EXT-X-ENDLIST tag marks a finished presentation, and EXT-X-INDEPENDENT-SEGMENTS promises that every segment starts on a keyframe.
#EXTM3U
#EXT-X-VERSION:3
#EXT-X-TARGETDURATION:6
#EXT-X-MEDIA-SEQUENCE:0
#EXT-X-PLAYLIST-TYPE:VOD
#EXT-X-INDEPENDENT-SEGMENTS
#EXTINF:6.000000,
seg_000.ts
#EXTINF:6.000000,
seg_001.ts
#EXTINF:6.000000,
seg_002.ts
#EXTINF:3.250000,
seg_003.ts
#EXT-X-ENDLIST
# Relative URIs resolve against THIS playlist's URL (RFC 8216), so the playlist and its
# segments can sit in one folder and move between hosts without editing a line.
# #EXT-X-ENDLIST tells the player the presentation is complete: no reloads needed.
Both playlists are small text files, so cache them too, but briefly if they can change. The relative URIs are not decoration: RFC 8216 says a relative URI is resolved against the playlist's own URI, which makes a whole folder relocatable between buckets and CDN hostnames.
How does the player choose and switch quality?
Adaptive bitrate is a client-side loop, and the server takes no part in it. After every segment the player measures how fast it arrived, smooths that into a throughput estimate, and picks the highest rendition whose declared BANDWIDTH fits under it with a safety margin. A second rule watches the buffer: when too little video is buffered it drops a rung immediately, because a stall is worse than a blurry picture. Real players such as hls.js and Shaka are more elaborate, but this is the core.
interface Rung { bandwidth: number; url: string } // BANDWIDTH from the master playlist
const SAFETY = 0.8; // assumed: use 80% of measured throughput, leave headroom
const PANIC_BUFFER_S = 5; // assumed: below this many seconds buffered, drop one rung now
let estimate = 0; // bits per second, smoothed
function onSegmentDownloaded(bytes: number, ms: number) {
const sample = (bytes * 8) / (ms / 1000);
// Exponentially weighted average: one slow segment nudges, not whipsaws, the estimate.
estimate = estimate === 0 ? sample : 0.7 * estimate + 0.3 * sample;
}
function pickRung(rungs: Rung[], current: number, bufferS: number): number {
// rungs are sorted ascending by bandwidth. Return the index to fetch the NEXT segment from.
if (bufferS < PANIC_BUFFER_S) return Math.max(0, current - 1);
let best = 0;
for (let i = 0; i < rungs.length; i++) {
if (rungs[i].bandwidth <= estimate * SAFETY) best = i;
}
return best;
}
// Worked example: measured 4,000,000 bps -> budget 3,200,000 bps.
// 1080p declares 5,478,000 (no), 720p declares 3,124,000 (yes) -> pick 720p.
// Switching only ever happens between segments, which is why they start on keyframes.
Switching only happens between segments, which is why the ladder cuts at identical times and why segments start on keyframes. A shorter segment lets the player react faster but costs more requests: at 10,000 viewers, 2-second segments mean 5,000 requests per second against 1,667 for 6-second ones. The constants in the code are assumptions, so tune them against real playback data rather than trusting mine.
How do signed URLs and DRM protect the video?
There are two different protections, and people mix them up. Access control decides who may download the segments at all: the CDN checks a signature or cookie with an expiry, and rejects everything else. Encryption and DRM decide what a downloader can do with the bytes. HLS supports AES-128 encryption of segments through the EXT-X-KEY tag, and commercial DRM systems add licence servers and hardware-backed key handling on top. The CDN documentation, CloudFront's for instance, distinguishes signed URLs for single files from signed cookies for many files.
// NestJS-style sketch: a short-lived token that covers a whole folder, not one file.
import { createHmac } from "node:crypto";
function playbackToken(videoId: string, userId: string, ttlSeconds = 4 * 3600): string {
const expires = Math.floor(Date.now() / 1000) + ttlSeconds;
const scope = "/v/" + videoId + "/"; // every rendition + segment below this path
const payload = scope + "|" + userId + "|" + expires;
const sig = createHmac("sha256", process.env.CDN_SIGNING_KEY!).update(payload).digest("hex");
return Buffer.from(payload + "|" + sig).toString("base64url");
}
// Wrong: sign each URL with a query string. The master playlist lists RELATIVE segment URIs,
// and a relative URI does not inherit the playlist's query string, so segments 403.
// Right: put the token in a cookie or a path prefix that every request under /v/<id>/ carries.
// Each CDN defines its own token format and validation; this only shows the shape of the idea.
A playlist references dozens of segments, so per-URL signing is awkward. A token covering the whole folder, in a cookie or the path, is the shape that works. Start with signed access and short expiries. Add DRM only when a contract or a rights holder requires it, because it brings licence servers, per-platform integration and real cost, and a determined viewer can still screen-record.
How is live streaming different from VOD?
Live uses the same segments and playlists, but the playlist keeps growing and the encoder never finishes. The differences are in timing and in what you can cache:
Aspect
VOD
Live
Media playlist
Complete and fixed, ends with EXT-X-ENDLIST.
No ENDLIST while live; the player keeps reloading it for new segments.
Encoding
Offline job, as slow as it needs to be.
Real time: each rendition must encode faster than the video plays.
Playlist caching
Long TTL is safe, it never changes.
TTL must be roughly one segment or less, or viewers see stale playlists.
Delay behind real time
None.
At least a few segment durations, derived below.
RFC 8216 sets the arithmetic. A server must publish a new playlist version no earlier than half a target duration and no later than 1.5 target durations after the previous one, and it must not shrink a live playlist below three times the target duration. With 6-second segments that is at least 18 seconds of playlist, so a viewer is typically 18 seconds or more behind the camera. Cutting segments to 2 seconds brings that near 6 seconds at the price of three times the requests. Lower-latency HLS variants exist, but they are a separate, more demanding design.
Which design should a small team choose? A checklist
Most products do not need to build this. If you stream a few hours of training video to a few hundred people, a managed video service is the sound choice. If you do build it, this is the order I would follow:
Keep the original upload in object storage and never overwrite it.
Run transcoding as a queued worker job, one input producing the whole ladder, with a keyframe interval equal to the segment length.
Publish HLS first, with fixed segment lengths, and set BANDWIDTH and CODECS from what the encoder really produced.
Put object storage behind a CDN and plan for a 99 percent hit ratio, not 95, using your own peak viewers times the highest rung's bitrate.
Protect access with expiring tokens on a path or cookie before considering DRM.
Add live only when you need it: it changes encoding, caching TTLs and delay all at once.
On a single VPS I would use the box for the API and the transcode worker, and keep every viewer-facing byte on object storage and a CDN. The VPS cannot serve the egress, as the arithmetic showed, but it can comfortably orchestrate the jobs. For the storage side, see the earlier comparison of cloud storage and MinIO.
Video streaming is a file pipeline with a smart client. Encode once into a ladder, cut every rung at the same keyframes, publish the segments as ordinary cacheable objects, and let the CDN absorb the egress while the player picks its own quality. Do the egress arithmetic first, because it, not storage, decides whether the design fits on your hardware.