Your API Works Fine at 100 RPS… What Actually Breaks at 10,000 RPS?
A practical guide to the CPU, Node.js event loop, database pools, Redis, queues, and backpressure failures that appear when an API grows from 100 to 10,000 RPS.
8 min read
The Same API, One Hundred Times the Traffic
At 100 requests per second, the API looks healthy:
p95 latency: 80 ms
Node.js CPU: 22%
DB connections: 12 / 50
Error rate: 0.1%Then a campaign sends 10,000 requests per second:
p95 latency: 8.4 seconds
Node.js CPU: 96%
DB connections: 50 / 50
Queue depth: 180,000
Error rate: 31%The route did not suddenly forget how to work. The system crossed several capacity boundaries at once.
Scaling is not one question—“Can the server handle 10K RPS?”—but a chain of limits:
Load balancer
|
v
Node.js instances
|
+--> PostgreSQL pool
+--> Redis
+--> downstream APIs
+--> queueThe first saturated dependency increases latency. Higher latency increases concurrency. Higher concurrency consumes more sockets, memory, and pool slots. A local bottleneck becomes a system-wide failure.
Start With a Traffic Budget
Suppose the endpoint does this for every request:
1 PostgreSQL read
1 Redis read
1 JSON serialization
1 call to payment-risk serviceAt 10,000 RPS, the downstream demand becomes:
PostgreSQL reads: 10,000 / sec
Redis operations: 10,000 / sec
Risk-service requests: 10,000 / secIf one response triggers five database queries, it becomes 50,000 queries per second. RPS is only the entry metric; fan-out determines the real workload.
Bottleneck 1: CPU and the Node.js Event Loop
Node.js handles I/O concurrency well, but JavaScript execution for a process still runs on an event loop. CPU-heavy work blocks unrelated requests.
app.post("/reports", async (req, res) => {
const rows = await loadRows(req.body.accountId);
// Large synchronous transformation blocks the event loop.
const report = buildAndCompressReport(rows);
res.json(report);
});At low traffic, a 30 ms synchronous task may go unnoticed. At high concurrency, those blocks form a queue inside the process.
Measure:
- CPU per instance
- Event-loop delay
- Event-loop utilization
- Garbage-collection pauses
- Heap growth
- Request duration by route
Move expensive CPU work to worker threads, a background job, or a separate service when it does not belong in the synchronous request path.
Bottleneck 2: The Database Connection Pool
Assume each application instance has a pool of 50 connections.
20 app instances × 50 connections = 1,000 possible DB connectionsIf PostgreSQL safely supports only a fraction of that concurrent workload, autoscaling the API makes the database problem worse.
Pool exhaustion usually looks like this:
Request arrives
|
v
Wait for DB connection: 1,800 ms
|
v
Query runs: 12 msThe query is fast, but the request is slow because it waits for admission.
Track pool metrics separately:
pool_active
pool_idle
pool_waiting
pool_acquire_duration_ms
query_duration_msIncreasing the pool is not automatically a fix. It can replace application waiting with database contention. First reduce unnecessary queries, optimize slow access patterns, and set a total connection budget across all instances.
That budget has to include every application replica. The deeper guide to database connection pooling in production shows how per-process pools multiply across a fleet.
Bottleneck 3: Cache Behavior Under Load
A 95% hit rate sounds excellent until 10,000 RPS produces 500 misses per second. If a hot key expires, those misses can arrive together.
The detailed guide on why a fast Redis cache can still destroy the database covers stampedes, TTL jitter, hot keys, and invalidation.
For capacity planning, test at least three states:
- Warm cache
- Cold cache
- Coordinated expiration of popular keys
Only testing a warm cache measures the easiest version of the system.
Bottleneck 4: Unbounded Concurrency
This code launches every downstream operation at once:
const results = await Promise.all(
customerIds.map((id) => calculateCustomerSummary(id)),
);For 20 IDs, it may be fine. For 50,000 IDs, it can overwhelm the database and downstream services.
Use bounded concurrency:
import pLimit from "p-limit";
const limit = pLimit(20);
const results = await Promise.all(
customerIds.map((id) => limit(() => calculateCustomerSummary(id))),
);The exact limit must come from measurement. The principle is that callers should not be able to create unlimited downstream work.
Bottleneck 5: Dependencies and Fan-Out
If one request calls four services and each is 99.9% available, the combined request path is less available than any one dependency.
API
|
+--> User service
+--> Inventory service
+--> Pricing service
+--> Recommendation serviceAt higher traffic, slow dependencies also hold more open requests. Set explicit timeouts, propagate cancellation, and decide which data is optional.
const response = await fetch(url, {
signal: AbortSignal.timeout(300),
});A timeout should protect a latency budget. It should not be immediately followed by uncontrolled retries; that can create a retry storm.
Bottleneck 6: Queues Without Backpressure
Moving work to a queue protects request latency only if producers are controlled and consumers have enough capacity.
API producers: 10,000 jobs/sec
Worker capacity: 6,000 jobs/sec
Backlog growth: 4,000 jobs/secAfter one minute:
240,000 queued jobsThe API appears successful while the promised work is hours behind.
Monitor:
- Queue depth
- Oldest job age
- Producer rate
- Consumer throughput
- Retry rate
- Dead-letter volume
Backpressure may mean returning 429, delaying admission, reducing optional work, or shedding low-priority traffic.
The same producer-consumer imbalance appears in streams, promise concurrency, and worker queues; the backpressure guide develops those controls end to end.
Diagnose the First Saturated Resource
Use a controlled load test that increases traffic in stages:
100 RPS -> hold -> observe
500 RPS -> hold -> observe
1K RPS -> hold -> observe
2K RPS -> hold -> observe
5K RPS -> hold -> observe
10K RPS -> hold -> observeAt each stage record:
| Layer | Useful signals |
|---|---|
| API | RPS, p50/p95/p99, errors, in-flight requests |
| Node.js | CPU, heap, GC, event-loop delay |
| Database | pool wait, query latency, locks, CPU, I/O |
| Redis | hit rate, latency, evictions, hot keys |
| Queue | depth, oldest age, throughput, retries |
| Dependencies | latency, timeout rate, error rate |
Stop when latency bends sharply even if errors have not started. That inflection point often identifies the first capacity boundary before the system collapses.
A Safer High-Traffic Shape
CDN / edge cache
|
Load balancer
|
-----------------------
| | |
API 1 API 2 API 3
| | |
-----------+-----------
|
--------------------------------
| | |
Redis connection pool Queue
| | |
| PostgreSQL Workers
| |
+------ bounded work ----------+This diagram is not a prescription. It shows the controls that matter:
- Cache repeatable reads
- Bound database concurrency
- Queue deferrable work
- Limit producer rate
- Scale stateless instances horizontally
- Keep timeouts and cancellation explicit
Trade-offs
| Technique | Benefit | Cost |
|---|---|---|
| More API instances | More request capacity | More DB connections and coordination |
| Caching | Fewer expensive reads | Invalidation and staleness |
| Queues | Smooth bursts | Delayed results and retry complexity |
| Rate limiting | Protects capacity | Rejects or delays callers |
| Precomputation | Cheap reads | Freshness and storage cost |
| Read replicas | More read throughput | Replication lag |
The best design is rarely the one with the most components. It is the one that makes overload behavior explicit.
Production Best Practices
- Define latency and error budgets per endpoint.
- Measure downstream operations per request.
- Set a database connection budget across the fleet.
- Track pool wait time separately from query time.
- Monitor event-loop delay, not only Node.js CPU.
- Bound concurrency for fan-out and batch work.
- Use queues for deferrable work and monitor job age.
- Apply load shedding before every dependency is saturated.
- Test cache-cold and dependency-slow scenarios.
- Scale based on the actual constrained resource.
Conclusion
At 100 RPS, spare capacity hides inefficient queries, synchronous CPU work, generous pools, and unbounded fan-out. At 10,000 RPS, those decisions become queues.
The investigation is always concrete:
Where does work wait?
Which resource saturates first?
What downstream demand does one request create?
How does the system reject or defer excess work?Horizontal scaling helps only when the next dependency can absorb the additional concurrency. Reliable high-throughput systems scale capacity and control admission together.
References
Related
Written by
Faisal
Software engineer writing about backend systems, Node.js, system design, scalable applications, and modern web and mobile development.