0:01
Every major API has ray limits. Make too
0:03
many requests to GitHub, Stripe or AWS
0:06
and your requests get rejected. How do
0:08
this system work? Ray limiters control
0:11
how many requests a client can make to
0:13
an API in a given time window. They
0:16
protect systems from overload while
0:17
maintaining fair access for legitimate
0:20
users. Let's tackle the core design
0:22
challenges. Let's go over the
0:24
requirements first. The system should
0:26
limit incoming requests based on
0:28
configurable rules like 100 API requests
0:31
per minute per user. When limits are
0:34
exceeded, the system should reject
0:36
requests with HTTP 429 and include
0:39
helpful headers showing rate limit
0:41
remaining and reset time. The system
0:44
should introduce minimal latency
0:45
overhead, say under 3 millisecond P95
0:49
per check. The system should be highly
0:51
available and accessible by multiple
0:53
servers.
0:55
Now that we understand what we need to
0:57
build, let's start with the simplest
0:58
approach, fixed window counting. We
1:01
divide time into fixed windows like one
1:03
minute intervals. Each user gets a
1:06
counter that resets at the start of each
1:08
window. Here's how it works. A user is
1:11
allowed 100 requests per minute. At the
1:13
start of each minute, their counter
1:15
resets to zero. Each request increments
1:18
to counter. When they hit 100 requests,
1:20
we reject additional requests until the
1:23
next minutes begins. We need to store
1:25
these counters somewhere fast. The
1:27
database is too slow for this. Today's
1:29
video is sponsored by Warped, the best
1:31
way to code with AI agents. Too often,
1:34
agents write code that's almost right,
1:36
leaving developers stuck debugging
1:38
instead of shipping. Warp is different.
1:41
Rank top of terminal bench and bench
1:43
verified. Warps agent understands your
1:46
context and writes production ready code
1:48
out of the box. Prom, review, and refine
1:51
all in one interface. No context
1:53
switching, no wasted time. You stay in
1:56
control and it pays off. On average,
1:58
users are saving over an hour a day with
2:00
Warp. Download Warp by clicking the link
2:03
in the description.
2:05
We're adding a database query to every
2:07
request, which could overload the very
2:09
system we're trying to protect. What
2:11
about in-memory storage and this server?
2:14
This would be very fast, but it only
2:16
works for a single server. When we scale
2:18
to multiple servers, each server would
2:21
have its own separate counters. A user
2:23
could make 100 requests to server A and
2:25
100 requests to server B, effectively
2:28
getting 200 requests per minute instead
2:30
of 100. Reddit solves both problems. Is
2:34
an in-memory data store that's shared
2:36
across all our servers. Reddis provides
2:38
primitives to increment counters and
2:40
reset them automatically. But fix window
2:43
have a critical flaw. Consider this
2:45
scenario. A user makes 100 requests in
2:48
the last 10 seconds of a minute, then
2:50
100 more requests in the first 10
2:52
seconds of the next minute. Both bursts
2:54
are within the 100 requests per minute
2:56
limit individually, but they've made 200
2:59
requests in just 20 seconds, which
3:01
clearly violates the intended ray limit.
3:04
This edge case happens at every window
3:06
boundary. We could use a better
3:08
algorithm.
3:09
The token bucket algorithm solves the
3:12
fixed window problem. It's the industry
3:14
standard used by companies like AWS and
3:16
Stripe. Think of it like this. Imagine a
3:19
bucket that hosts tokens. New tokens are
3:22
added at a steady rate. Each request
3:24
consume one token. When there are no
3:26
tokens left, we reject the request.
3:29
Let's see how this fixes our window
3:31
boundary problem. We set a bucket
3:33
capacity of 100 tokens that refills at
3:36
100 tokens per minute. During quiet
3:38
periods, tokens accumulate. When user
3:41
makes burst requests across window
3:43
boundaries, they consume accumulated
3:45
tokens but can't exceed the refill rate
3:48
over time. The key insight is that
3:50
tokens accumulate during quiet period.
3:52
This allow legitimate traffic bursts
3:55
while maintaining the overall rate
3:56
limit. A user can game the system by
3:59
timing the request to window boundaries.
4:01
The algorithm uses two parameters.
4:03
Bucket capacity determines burst size
4:06
and refill rate determines sustained
4:08
throughput. A capacity of 100 with a
4:11
refill rate of 100 per minute allows up
4:13
to 100 requests instantly that maintains
4:16
exactly 100 requests per minute
4:18
long-term. There are other algorithms
4:20
like sliding window logs, sliding window
4:22
counters, and leaking buckets that solve
4:25
the fixed window problem differently.
4:27
But token buckets strike the best
4:29
balance between simplicity and
4:31
effectiveness for most use cases. Now
4:33
that we have our algorithm, we need to
4:35
decide where to implement it in our
4:37
system architecture. We have three
4:39
options for where to place our ray
4:41
limiter. Client side ray limiting puts
4:43
the logic in client applications, but we
4:46
can't trust clients to enforce their own
4:48
limits. Malicious users can modify the
4:50
code or bypass restrictions entirely.
4:53
Serverside ray limiting embeds the logic
4:56
in our application code. This gives us
4:58
complete control over the algorithm and
5:00
keeps everything in one place. The
5:02
downsides is that ray limiting gets
5:04
mixed with business logic and each
5:06
service needs its own implementation.
5:09
Middleware ray limiting uses a dedicated
5:11
service between clients and the APIs.
5:14
This could be an API gateway, a reverse
5:16
proxy or a custom service. This keeps
5:19
ray limiting separate from business
5:21
logic and provides a single place to
5:23
manage policies. The trade-off is an
5:26
increase in system complexity. Which
5:28
should we choose? It depends on our
5:30
situation. If we already have an API
5:33
gateway handling authentication, adding
5:35
rate limiting there makes sense. If we
5:37
need a custom algorithm, server side
5:40
gives us flexibility. For most systems,
5:42
middleware offers the best balance of
5:44
control and operational simplicity. Our
5:47
system architecture looks like this. We
5:49
store ray limiting rules in a
5:51
configuration service. These rules
5:53
define limits like premium users get a
5:56
th00and requests per hour or free users
5:59
get 100 requests per hour. The
6:01
middleware fetches these rules and
6:02
stores token bucket state in Reddus.
6:05
When a request arrives, the middleware
6:07
identifies the user, fetches their token
6:09
bucket from Reddus and checks if tokens
6:11
are available. If yes, it decrements the
6:14
token count and forwards the request to
6:16
API servers. If no token remains, it
6:19
returns HTTP 429 with header showing
6:22
when the user can retry. This works
6:25
great for a single rail limited server.
6:27
But what happens when we need to scale
6:28
to multiple servers, we run into race
6:31
conditions. Here's what happens. Server
6:33
A reads a counter value of three from
6:35
radius. At the same time, server B also
6:37
reads three. Both servers check their
6:40
limits, decide the request is okay,
6:42
increment the value to four, and write
6:44
it back to Reddus. The counter now shows
6:47
four instead of the correct value of
6:48
five. We've lost a count. We can solve
6:51
this with atomic operations. Reddus
6:53
supports these two lure scripts that
6:55
bundle the read, check, and increment
6:57
into a single indivisible unit. This
7:00
prevents race condition entirely. We've
7:03
covered the essential building blocks
7:04
for ray limiters in this video. There
7:06
are other interesting topics we didn't
7:08
cover. How do we handle geographic
7:10
latency with multi-reion deployment? How
7:13
do we handle hot keys when a few users
7:15
generate most traffic? How do we upload
7:18
rate limiting rules without restarting
7:20
servers? Should the system fail open or
7:22
fail closed when key components go down?
7:25
These are all interesting topics worth
7:27
exploring.
7:29
Ready to ace your next technical
7:30
interview? Join our community where we
7:32
offer comprehensive courses on system
7:34
design, coding, behavioral questions,
7:38
machine learning, and object-oriented
7:40
design. Learn more at byitebico.com.