What happened
At 14:02 UTC we began a canary release of checkout-api v4.212.0, which moved tax calculation from a nightly cached rate table to a live quote from our tax vendor. The canary ran clean at 5% of traffic for 17 minutes and was promoted to 100% at 14:19.
The new code called the tax vendor from inside an open Postgres transaction. Under low traffic the extra latency was invisible. At full traffic — roughly 340 checkout requests per second at the Thursday afternoon peak — the average connection hold time rose from 12 ms to 140 ms. The pgbouncer pool has 200 connections; at that hold time, demand exceeded capacity. Requests began queueing for connections, queue wait reached 1.9 s, and the 10-second edge timeout started firing.
We rolled back at 14:52. Latency recovered within three minutes of the rollback reaching 100%, and the error rate returned to baseline at 15:13.
Timeline
-
14:02UTCCanary deploy begins
checkout-api v4.212.0goes to 5% of traffic. Release checks pass: error rate flat, p99 at 700 ms. -
14:19Promoted to 100%
Automated canary analysis reports no regression. Rollout completes across all 42 pods by 14:24.
-
14:26Degradation starts
Checkout p99 crosses 2 s. No alert fires — the page threshold is p99 above 4 s sustained for 5 minutes.
-
14:31First customer report
A support engineer posts two merchant reports of hanging checkouts in
#support-escalations. -
14:34Page fires · acknowledged in 2 min
checkout_p99_highpages the on-call. S. Okafor acknowledges at 14:36 and opens an investigation. -
14:38Incident 2041 declared (SEV-3)
Incident channel and status page draft are created automatically. Two responders join.
-
14:41Pool saturation identified
pgbouncer shows 200/200 connections active and 1.9 s average client wait. Query time itself is normal at 4 ms p99.
-
14:47Raised to SEV-2
Error rate crosses 5%. Responders correlate onset with the 14:19 promotion and choose rollback over a forward fix.
-
14:52Rollback begins
One command reverts to
v4.211.4. Pods drain and replace in waves of eight. -
15:01Rollback at 100% · recovery
p99 falls to 780 ms within three minutes. Connection wait returns to under 5 ms.
-
15:13Error rate at baseline
Checkout errors back to 0.04%. Degradation window closes at 47 minutes.
-
15:40Cart recovery sweep
1,204 shoppers had already retried successfully. Recovery emails go to the remaining 676 preserved carts.
-
16:20Incident resolved
Sixty minutes of stable metrics. Status page updated and the incident is closed.
Root cause
An external HTTP call was placed inside an open database transaction
v4.212.0 replaced the cached tax rate lookup with a live call to the tax vendor. The call was written inside the existing BEGIN … COMMIT block that reserves the cart, writes the order draft, and decrements inventory. The transaction therefore held a pooled Postgres connection for the full duration of a network round trip.
The vendor is not slow by its own standards: median quote latency was 90 ms and p99 was 1.4 s during the incident, both inside its published SLA. The defect is ours. Holding a connection across any network call converts vendor latency into database pool pressure.
Pool exhaustion is not linear. Below the saturation point, higher hold time is absorbed with no visible effect; above it, wait time climbs steeply. At 340 RPS and 140 ms hold time the theoretical demand was roughly 48 connection-seconds per second against a 200-connection pool — comfortable on paper, but request arrival is bursty, and the p99 tail of the vendor call pushed concurrent demand past 200 for several seconds at a time. Each excursion queued requests that were still queued when the next arrived.
Why the canary did not catch it. At 5% of traffic the service handled about 17 RPS. Connection demand never approached the pool limit, so the regression in hold time produced no change in any signal the canary analysis watches. The canary measured the symptom we expected, not the resource we consumed.
Contributing factors
-
01
The canary window was short and the traffic unrepresentative
Seventeen minutes at 5% is enough to catch an error-rate regression and nothing else. Saturation failures need a traffic share near the knee of the curve to appear at all.
-
02
We had no metric for connection hold time
Dashboards showed query duration, which stayed flat at 4 ms p99 throughout. The quantity that actually changed — how long a request held a connection — was not measured anywhere.
-
03
The alert threshold sat above the pain threshold
Checkout became unusable for many shoppers at 2 s p99. The page fires at 4 s for 5 minutes. Customers reported the problem three minutes before our monitoring did.
-
04
The vendor client had no timeout budget
The SDK default timeout is 10 s — longer than the entire checkout request budget. A slow quote could hold a connection for ten seconds with no circuit breaker to stop it.
-
05
The load test used a 5 ms mock of the vendor
Staging soak tests passed at 400 RPS against a fake that answered instantly, which removed the exact property that caused the outage.
What went well · what was luck
Went well
- Rollback was a single command and took nine minutes end to end.
- Responders chose to roll back rather than debug forward, before understanding the cause.
- Cart state was durable, so 64% of affected shoppers recovered on their own retry.
- Incident tooling created the channel, timeline, and status page draft without manual work.
Was luck
- The release landed on a Thursday afternoon, not during a weekend promotion, when peak traffic is 4× higher.
- A support engineer happened to be reading merchant reports in real time.
- The vendor's p99 was 1.4 s. At its SLA ceiling of 3 s the pool would have collapsed within four minutes of full rollout.
Action items
| ID | Action | Owner | Priority | Due | Status |
|---|---|---|---|---|---|
| AI-2041-1 | Move every external call out of the checkout transaction. Quote tax before BEGIN and pass the result in. |
P. RamanPayments Platform | P0 | 2025-08-20 | In progress |
| AI-2041-2 | Emit connection hold time and pool wait as first-class metrics. Page at 250 ms p99 hold time for 3 minutes. | S. OkaforObservability | P0 | 2025-08-22 | Done |
| AI-2041-3 | Set a 400 ms timeout and a circuit breaker on the tax vendor client, falling back to the cached rate table. | D. IversenPayments Platform | P1 | 2025-08-27 | In progress |
| AI-2041-4 | Lower the checkout latency page to p99 above 1.5 s for 3 minutes, and add a shopper-facing checkout success-rate SLO. | S. OkaforObservability | P1 | 2025-08-29 | Done |
| AI-2041-5 | Change canary policy for checkout-path releases: 5% for 30 minutes, then 25% for 30 minutes, with pool metrics in the analysis. | M. BellRelease Engineering | P1 | 2025-09-03 | Not started |
| AI-2041-6 | Replace vendor mocks in the staging soak with a latency-injecting fake set to the vendor's real p99 of 1.4 s. | P. RamanPayments Platform | P2 | 2025-09-10 | Not started |
| AI-2041-7 | Add a pool-saturation section to the checkout runbook, including the pgbouncer queries used at 14:41. | D. IversenPayments Platform | P2 | 2025-09-05 | Done |
Note on blame
The change was reviewed by two engineers and passed every gate we had. Nobody involved could have seen this from the diff, because the failure lives in the interaction between a transaction boundary, a vendor's tail latency, and a pool size set four years ago. The fix is not more care during review. It is a metric that makes connection hold time visible, a canary that runs long enough to saturate something, and a rule that keeps network calls out of transactions.