Devika Rao · Bengaluru

Tail Latency

Notes on distributed systems, storage engines, and the incidents that taught me more than the design docs did. Roughly one post a month, always with the numbers.

A p99 that only got worse when we added capacity

By Devika Rao · 4 March 2026

We doubled the fleet and the p99 went up by 40ms. That sentence sat in our incident channel for two days before anyone could explain it, and the explanation turned out to be a good one, so here it is written down.

The setup

A read path: stateless API pods, a connection pool per pod, and a Postgres primary with two replicas. Reads go to replicas through a pooler. Median 4ms, p99 around 60ms at 12k reads/sec, which we were happy with.

Traffic forecasts said 20k/sec by April, so we went from 40 pods to 80. Median stayed flat. p99 went from 60ms to 101ms within an hour of the rollout, and stayed there.

What we checked first, and why it was all wrong

Every dashboard we owned said nothing had changed. The thing that had changed was not on a dashboard: the number of connections.

The actual cause

Each pod opened a pool of 20 connections at startup and kept them warm. At 40 pods that is 800 connections into the pooler; at 80 pods it is 1,600, against a pooler configured for 900 server-side slots. Past that ceiling the pooler queues, and its queue is FIFO with no fairness across clients.

So the median request still found a free slot immediately — plenty were free most of the time. But the unlucky tail now waited behind a queue that had not existed at all a week earlier. More capacity in the stateless tier had made the stateful tier’s contention worse. Obvious in hindsight, invisible in the metrics we had.

Adding stateless capacity is only free until it multiplies your connection count against something that is not stateless.

The fix, in order of how much it helped

Three months on, p99 at 19k reads/sec is 54ms on 80 pods. The number that mattered was never CPU. It was a multiplication we were not doing.

Everything else

New posts by email

One email per post, roughly monthly. No digests, no course upsell, no “hey quick question” follow-ups.

Connect a runtime to start collecting subscribers.