How Request Bottlenecks and Cascading Failures Work
Understanding system limits before they fail in production.
When a web service receives a request, it usually hands it off to an application server which may then query a database. If any part of the system receives more requests than it can handle, it becomes a bottleneck. The excess requests increase latency, and eventually, the system starts dropping requests.
The Anatomy of a Bottleneck
Imagine a system with 2 application servers. Under normal conditions, they can handle the traffic perfectly. If the Request Rate spikes, the application servers will max out their capacity. When the application layer is the bottleneck, scaling the database won't help because requests are being dropped before they even reach the database.
If you scale the application servers without caching, the bottleneck often shifts down to the database, causing DB Load (modeled in requests per second) to saturate its available capacity.
Cascading Failures
In a cascading failure, one degraded component brings down the rest of the system. For example, if your Redis cache goes offline, all read traffic goes directly to the database. The database quickly becomes overloaded, causing requests to back up in the application servers, increasing Latency and Error Rate until the entire application is effectively offline.
Interactive Experiment
The best way to understand this is to break it yourself. In the Request Flow Simulator, try the following:
- Baseline: Run the simulation at a low Request Rate. Observe the stable Latency and Throughput.
- Spike: Increase the Request Rate dramatically. Watch the App Servers saturate and Error Rate skyrocket. The bottleneck is the "app".
- Shift the Bottleneck: Increase the App Servers. The app bottleneck is resolved, but now the Database might become the new bottleneck as DB Load spikes.
- Mitigate: Increase the Cache Hit Rate to shield the database and clear the bottleneck.