0:01
WhatsApp and Facebook Messenger process
0:03
billions of messages daily. Behind every
0:05
message is a distributed system
0:07
delivering text instantly across the
0:09
globe. How do we design a test system
0:11
that can handle real-time messaging at
0:13
this scale? Let's break it down. There's
0:16
a lot to cover when designing a test
0:18
system, so let's focus on the core
0:20
features. Our test system supports
0:22
one-on-one and group chats up to 100
0:24
participants. We guarantee message
0:26
delivery. We store messages for offline
0:29
users for a limited time. This is
0:31
different from systems that keep all
0:33
messages in central storage forever. We
0:36
also show online presence with that
0:38
familiar green dot. We're designing for
0:41
a growing chat app that starts with
0:43
thousands of users and scales to
0:45
millions.
0:46
How do clients and servers communicate
0:48
in a chat system? We could use HTTP for
0:51
everything. When someone sends a
0:53
message, the client makes an HTTP post
0:56
request. The server acknowledges and
0:58
forwards the message to the recipient.
1:00
This works well for sending, but
1:02
receiving is problematic. HTTP is client
1:05
initiated. Servers can easily push
1:08
messages directly to clients. How do we
1:11
get messages to users instantly? Polling
1:13
is an option. The client repeatedly ask
1:16
any new messages. Most requests return
1:19
empty. This waste server resources. Long
1:22
pollings hold connections open longer
1:24
until messages arrive or the connection
1:27
times out. This reduces requests, but we
1:30
can do better. Websocket offers a better
1:32
solution. After an HTTP handshake, the
1:35
connection upgrades to a persistent
1:37
channel. The server can push messages
1:39
instantly. However, websocket
1:42
connections are more difficult to scale,
1:44
which we'll address later in the video.
1:46
We could use websocket for both
1:48
directions, but we keep things simple.
1:50
HTTP works fine for sending and scales
1:53
easily. Websocket handles the real
1:55
challenge receiving messages instantly.
1:58
So we choose a hybrid approach. HTTP for
2:01
sending messages and websocket for
2:03
receiving them. Now we have chosen our
2:05
protocols. Let's design the overall
2:07
system. We need multiple components
2:10
working together to handle a large
2:12
number of users. We divide the system
2:14
into three layers. Salus services,
2:16
staple services and third party
2:18
integrations. Sailor services handle
2:21
authentication, user profiles, and
2:23
message sending through REST APIs. Load
2:26
balances distribute requests across
2:28
multiple servers. This is a known design
2:30
pattern that scales easily by adding
2:32
more servers as needed. Chat servers
2:35
form the stable core. Each client
2:37
maintains a persistent websocket
2:39
connection to one chat server. Modern
2:42
servers can handle tens of thousands or
2:44
hundreds of thousands of concurrence
2:46
websocket connections. This capacity
2:49
determines how many servers we need as
2:51
we scale. When someone sends you a
2:53
message, but you are not connected to a
2:55
chat server, the system needs another
2:57
way to reach you. Third party services
2:59
like Apple's APNS or Google's FCM
3:03
deliver push notifications when users
3:05
are offline. They maintain persistent
3:07
connections to millions of devices. This
3:10
is not something we want to reinvent.
3:12
Our architecture handles online users
3:15
well, but what about offline scenarios?
3:18
When someone sends you a message while
3:20
your phone is off or when network
3:22
connections fail during transmission,
3:24
how do we guarantee delivery? The inbox
3:27
pattern solves this. Every users gets an
3:29
inbox, a personal queue that stores and
3:32
deliver messages. When someone sends you
3:34
a message while you're offline, it waits
3:36
in your inbox for a limited time. Here's
3:38
how it works. When someone sends you a
3:41
message, the server tries to deliver it
3:42
immediately if you're online. Your
3:45
device gets the message through
3:46
websocket and sends back an
3:48
acknowledgement. Done. No storage
3:50
needed. But if you're offline or the
3:52
delivery fails, the message goes into
3:55
your inbox. It waits there until you
3:57
come back online. The acknowledgement
3:59
mechanism provides reliability for inbox
4:01
messages. When your device receives a
4:04
message from the inbox, it sends back an
4:06
act. The server only removes the message
4:08
from the inbox after getting this
4:10
acknowledgement. No act means the server
4:12
will retry later. When you come back
4:14
online, your chat server checks your
4:16
inbox for waiting messages. All
4:19
accumulated messages download to your
4:21
device in order, bringing you up to
4:23
date. The inbox patterns works great,
4:26
but now we have a new challenge. With
4:28
multiple chat servers handling different
4:30
users, how do messages travel between
4:32
them? We need a way to route messages
4:34
between chat servers when users are
4:36
connected to different servers. The most
4:38
scalable approach uses a combination of
4:41
service discovery and direct server
4:43
communication. Let's say Alice sends a
4:46
message to Bob. Alice's chat server
4:48
first checks if Bob is online. The
4:50
server queries a user presence service
4:52
that tracks which chat server Bob is
4:54
connected to. This service acts like a
4:57
directory mapping users to the current
4:59
server locations. If Bob is online,
5:02
Alice's server makes a direct RPC call
5:04
to Bob's chat server. Bob server gets
5:07
the message and immediately pushes it
5:09
through Bob's websocket connection. This
5:11
direct approach minimizes latency. This
5:14
scales well with service discovery
5:16
systems
5:18
that help servers find each other. Chat
5:21
platforms like Discord uses direct
5:23
serverto-s server communication pattern.
5:25
If Bob is offline, the message goes to
5:27
his inbox for later delivery. And we
5:30
trigger a push notification through
5:32
third party services. One-on-one
5:34
messages works smoothly now, but group
5:36
chat introduce a new complexity. How do
5:39
we efficiently deliver one message to
5:41
100 different users? We use a fan out
5:44
pattern. When someone sends a group
5:46
message, the server checks which members
5:48
are online and make RPC calls to deliver
5:50
immediately to their devices. Offline
5:53
members get the message stored in the
5:55
inboxes. This works well for a
5:57
requirement of up to 100 members since
5:59
the fan out workload scales with group
6:01
size. With messaging sorted out, let's
6:04
work on another feature, online
6:06
presence. The green dot showing online
6:08
status seems simple but creates
6:10
interesting technical challenges. We use
6:13
heartbeats to track presence. We can
6:15
rely on websocket connection stay alone.
6:18
connection can appear alive even when
6:20
the app is backgrounded or the device is
6:22
sleeping. Clients send a websocket ping
6:25
frame every 30 seconds to their chat
6:27
server. The server tracks the timestamp
6:29
of the last ping it receives for each
6:31
user. If the chat server gets no ping
6:34
for 60 seconds, it marks the user
6:36
offline. This approach smooths out brief
6:38
disconnection in tunnels, elevators, or
6:41
areas with poor coverage. We built a
6:44
working chat system, but success brings
6:46
new problems. As our user base grows
6:49
from thousands to millions, which
6:51
components bricks first connections
6:54
likely becomes the first bottleneck. As
6:56
user count grows, we need more chat
6:58
servers to handle the connections. The
7:01
database can also hit limits as we
7:03
scale. As millions of users go offline
7:05
and come back online, the inboxes filled
7:08
up and empty constantly. We can split
7:10
users across multiple database servers
7:12
based on their user ids. For global
7:15
reach, we deploy chat servers in
7:17
multiple regions. Users connect to their
7:19
nearest region for low latency, but
7:21
cross region message routing as
7:23
complexity.
7:25
This covers the core chat system
7:27
architecture. There are other areas
7:29
worth considering. What about message
7:31
ordering? When messages travel across
7:33
distributed servers, they can arrive out
7:35
of order. Message IDs, timestamps, and
7:38
vector clocks help solve this. Security
7:41
and encryption add another layer of
7:43
complexity. End-to-end encryption
7:45
requires careful key exchange
7:46
mechanisms. Group chess makes this even
7:49
more challenging. Media handling is its
7:52
own beast. Sharing images and videos
7:54
requires compression, CDNs, and
7:57
progressive loading to keep things fast.
7:59
Then there features like red receipts
8:01
and typing indicators. Broadcasting
8:03
typing status to group members sounds
8:05
simple but can easily overwhelm your
8:08
system without careful design. One more,
8:10
we need ray limiting and abuse
8:12
prevention. Stopping spams and API abuse
8:15
while keeping the experience smooth is a
8:17
delicate balance. We are barely
8:20
scratching the surface, but this is how
8:22
we design the core features of a chat
8:24
system that handles real-time messaging
8:26
at scale.
8:28
Ready to ace your next technical
8:30
interview? Join our community where we
8:32
offer comprehensive courses on system
8:34
design, coding, behavioral questions,
8:37
machine learning, and object-oriented
8:39
design. Learn more at bitebico.com.