Interactive / No signup / 3 minutes

Production is on fire.
You’re on call.

No lectures. No setup. Three engineering puzzles. Break things, fix the bottleneck, and see why your solution works.

● TRAFFIC SURGE / CASE 01

650 people per second. Your servers can't keep up.

A product launch just made the front page. Requests are failing faster than your team can respond.

YOUR OBJECTIVE

Keep the site online without throwing unlimited servers at it.

Deliver ≥ 98% successful requests, ≤ 300ms response time, using ≤ 8 servers.

1

Start here: your app servers are dropping traffic.

Change a setting, watch the live numbers, then tap Run the fix to clear all three targets. Need help? Use the guided move below.

Successful requests

30.8%

200 of 650 req/s

Response time

195ms

Target: ≤ 300ms

Server budget

2 servers

Maximum: 8

THE CONTROL ROOM

Make your move

Each server accepts 100 requests/sec in this model.

More cache hits means fewer database requests.

Redis cache

LIVE REQUEST PATH

Incoming

650 requests/sec

↓

App servers

200 / 200 capacity

↓

Redis

50 served from cache

↓

Database

150 / 500 capacity

Failed requests/sec: 450

Live diagnosis

The app layer is dropping requests before they reach the cache. More app capacity is the first move.

Extra requests served / sec

0

Response time change

0ms

Recommended next move

Your app servers can handle only 200 of 650 incoming requests each second. Some requests never reach Redis.

Moves tried: 0 · This is a simplified deterministic learning model—not a real-world performance benchmark. Metrics update instantly as you change settings.