0:00
Upload video, watch video. Sounds
0:02
simple. YouTube processes 500 hours of
0:05
video every minute. That's 30,000 hours
0:08
of content uploaded every hour. The
0:11
engineering behind this interface is
0:13
complex. Let's design YouTube. Not the
0:16
billion user version, but a smaller
0:18
system that we can learn from. This
0:20
video covers the core components. We'll
0:22
focus on video upload and streaming.
0:24
There's much more to YouTube. search,
0:27
recommendations, comments, monetization.
0:30
Each could be its own deep dive. Today,
0:32
we'll build a foundation you can extend.
0:35
Here are the requirements we'll focus
0:36
on. People upload massive videos. A
0:39
10-minute 4K video ranges from 1.5 GB
0:42
highly compressed to 30 GB in ProRes.
0:46
When uploading huge videos, users expect
0:48
resumable uploads. For viewing, playback
0:51
must adapt to network condition in real
0:53
time. When your connection drops from 25
0:56
megabit per second to 2 megabit per
0:58
second, playback shouldn't stop. The
1:00
upload challenge reveals the first
1:02
design decision. We could route large
1:04
videos through our API servers,
1:06
configure streaming instead of
1:08
buffering, increase time out from 30
1:10
seconds to 30 minutes, handle partial
1:13
uploads is all solvable. But why make
1:16
our servers handle gigabytes of pass
1:17
through traffic when they could be
1:19
serving actual API requests? There's a
1:22
better pattern. use pre-signed URLs. Our
1:25
API server generates a temporary signed
1:27
URL that grants direct upload permission
1:30
to blob storage. The client uploads
1:32
straight to storage. Our server stay
1:35
free to handle other work. Blob storage
1:37
gives us another benefit. We get
1:39
multiport uploads built in. The client
1:42
splits the videos into chunks. Each
1:44
chunk is 5 to 10 megabyte. It's small
1:47
enough to upload quickly on slower
1:49
connections. Large enough to minimize
1:51
overhead. Each chunk get a SH 256
1:54
fingerprint. The client uploads chunks
1:56
in parallel. Six chunks uploading
1:58
simultaneously is common. This pattern
2:01
appears in Dropbox, Google Drive, and
2:03
backup services. Any system handling
2:06
large files uses similar approaches.
2:09
The processing pipeline presents a
2:11
different challenge. Video transcoding
2:13
burns massive compute cycles. User
2:16
uploads videos in many different
2:17
formats. iPhone record in HGVC. Android
2:21
phones uses H.264.
2:23
Someone uploads a 4K Pro file from Final
2:26
Cut Pro. We need these videos playable
2:28
on every device. Old Android phones run
2:31
ancient Android. Smart TVs from 2018,
2:35
web browsers that haven't updated in
2:36
years. The solution is a processing
2:39
pipeline that converts one video into
2:41
many versions. We generate multiple
2:43
resolutions from 2160p down to 240p. We
2:47
need so many because network conditions
2:49
and device speeds vary widely. Someone
2:52
on fiber with a fast device needs 4K.
2:55
Someone on 3G with an ancient phone
2:57
needs 240p. Then consider codecs. H264
3:02
works everywhere but uses more
3:03
bandwidth. VP9 saves bandwidth but older
3:06
devices can decode it. AV1 saves even
3:09
more bandwidth but needs powerful
3:11
hardware. We encode in all three. Next,
3:14
we package these in containers. MP4 for
3:17
maximum compatibility. Webb for web
3:20
optimization. Now one upload becomes 15
3:23
to 20 files. How do we process
3:25
efficiently? Model the workflow as a
3:28
DAG. DAG stands for the directed as
3:30
cyclic graph. Each processing step is a
3:33
note. Dependencies are edges. The
3:36
asyclic part matters. It ensures task
3:38
can complete without circular
3:40
dependencies. First split the video into
3:43
segments. Videos have key frames every
3:45
two to 10 seconds. These frames send
3:48
alone without referencing others. Split
3:50
a key frames now use segment processes
3:53
independently. The workflow splits into
3:55
multiple streams. Video, audio, and
3:58
metadata each take their own path
4:00
through the system. Video segments fan
4:03
out to hundreds of workers while one
4:05
machine transcodes segment one to 1080p.
4:08
Another handles segment 2 to 720p. Audio
4:11
processing runs in parallel on different
4:13
hardware. Thumbnail generation and
4:15
subtitle extraction happen on their own
4:17
dedicated workers. This is the power of
4:20
modeling work as a DAG. One video
4:22
becomes hundreds of parallel tasks
4:24
across a worker farm. As each task
4:26
completes, results flows to the next
4:28
stage. Sequential processing would take
4:31
hours. Parallel processing completes in
4:33
minutes. Now streaming. Modern video
4:36
streaming uses adaptive bitray
4:38
streaming. The video player doesn't
4:40
download one file. It downloads
4:42
segments, small chunks of videos, each a
4:44
few seconds long. When network bandwidth
4:46
is high, the player fetches 1080p
4:48
segment. When bandwidth drops, it
4:51
switches to 480p segments. The
4:53
transition is usually seamless. This
4:55
work through manifest files. The primary
4:57
manifest lists all available formats.
5:00
Each format then has its own media
5:02
manifest with URLs for every segment.
5:04
The player reads manifests, monitors
5:07
bandwidth using the download speed of
5:08
the recent segments, and fetches the
5:10
probate segments. All this happened
5:12
invisibly. The player makes HTTP range
5:15
requests. Give me 1,000 to 2,000 of
5:19
segment five of the 720p version. This
5:22
enables instance seeking to jump to
5:24
minute 47. The player calculates which
5:26
segments to fetch. These segments are
5:28
stores in CDN's, content delivery
5:31
networks. Popular videos cache across
5:33
edge servers worldwide. A viewer in
5:35
Tokyo get segments from Tokyo, not
5:38
California. Geographic proximity means
5:40
lower latency, better streaming. We've
5:43
just scratched the surface. Several
5:45
areas deserve deeper exploration. The
5:48
hot video problems challenges every
5:50
video platform. When one video goes
5:52
viral, millions request it
5:54
simultaneously. We need metadata caching
5:56
and database hotspot prevention. Cost
5:59
optimization requires trade-offs. Do we
6:01
transcode everything immediately or
6:04
popular formats first? Should rare
6:06
formats use ondemand transcoding? When
6:08
do we migrate to code storage? Pipeline
6:11
optimization can improve latency. Stop
6:13
processing segments as they arrive
6:15
instead of waiting for complete upload.
6:17
Pipeline to upload and processing for
6:20
faster availability. We could explore
6:22
geographic CDN placement, readwrite
6:25
ratios or lazy transcoding strategies.
6:28
Each topic could be its own deep dive,
6:30
but the fundamentals remain the same.
6:32
What makes this design work at scale?
6:34
Direct uploads keep servers free for
6:36
actual logic. DAX transforms sequential
6:38
bottlenecks into parallel workflow.
6:41
Adaptive streaming ensures smooth
6:42
playback regardless of network
6:44
conditions. Once we understand these
6:46
principles, we see them everywhere.
6:48
Large file sharing, machine learning
6:50
pipelines, and live streaming all build
6:52
on these same foundations.
6:55
Ready to ace your next technical
6:56
interview? Join our community where we
6:58
offer comprehensive courses on system
7:01
design, coding, behavioral questions,
7:04
machine learning, and object-oriented
7:06
design. Learn more at bitebico.com.