This video explains consistent hashing, a method for distributing requests among servers efficiently. It describes how requests and servers are mapped onto a circular 'hash ring' and how requests are assigned to the nearest server clockwise. The video also shows how consistent hashing minimizes changes in server load when servers are added or removed, and how using virtual servers (multiple hash functions) helps prevent uneven load distribution.

Key Takeaways

1

The core problem consistent hashing solves is the complete remapping of data when servers are added or removed, which happens with traditional load balancing.

2

In consistent hashing, both requests and servers are mapped to points on a circular 'hash ring' using a hash function.

3

When a request comes in, it is assigned to the nearest server found by moving clockwise around the hash ring from the request's point.

4

Adding or removing a server with consistent hashing only affects a small portion of the requests, minimizing the overall data movement and re-assignment compared to other methods.

5

A potential issue with consistent hashing with a small number of servers is that load can become unevenly distributed among servers.

6

To prevent skewed load distribution, 'virtual servers' are used by applying multiple hash functions to each server ID, creating multiple points for each physical server on the hash ring.

7

Having multiple virtual server points for each physical server increases the likelihood of an even load distribution and minimizes the impact of server additions or removals.

8

Consistent hashing is widely used in distributed systems for load balancing, such as in web caches and databases, because of its flexibility and efficiency.

What is CONSISTENT HASHING and Where is it used?

Gaurav Sen
Feedback