Production is on fire.
You’re on call.
No lectures. No setup. Three engineering puzzles. Break things, fix the bottleneck, and see why your solution works.
650 people per second. Your servers can't keep up.
A product launch just made the front page. Requests are failing faster than your team can respond.
YOUR OBJECTIVE
Keep the site online without throwing unlimited servers at it.
Deliver ≥ 98% successful requests, ≤ 300ms response time, using ≤ 8 servers.
Start here: your app servers are dropping traffic.
Change a setting, watch the live numbers, then tap Run the fix to clear all three targets. Need help? Use the guided move below.
Successful requests
30.8%
200 of 650 req/s
Response time
195ms
Target: ≤ 300ms
Server budget
2 servers
Maximum: 8
THE CONTROL ROOM
Make your move
Each server accepts 100 requests/sec in this model.
More cache hits means fewer database requests.
Redis cache
LIVE REQUEST PATH
Incoming
650 requests/sec
App servers
200 / 200 capacity
Redis
50 served from cache
Database
150 / 500 capacity
Failed requests/sec: 450
Live diagnosis
The app layer is dropping requests before they reach the cache. More app capacity is the first move.
Extra requests served / sec
0
Response time change
0ms
Recommended next move
Your app servers can handle only 200 of 650 incoming requests each second. Some requests never reach Redis.
Moves tried: 0 · This is a simplified deterministic learning model—not a real-world performance benchmark. Metrics update instantly as you change settings.