Aperture Commerce / Engineering
Postmortem · Doc INC-2041 · Rev 3
SEV-2 Resolved Checkout · Payments Platform

Incident 2041 — Checkout latency, August 14

A release added a synchronous tax-quote call inside the checkout database transaction. Connection hold time grew 11×, the pool saturated, and checkout degraded for 47 minutes before a rollback restored service.

Summary of impact All times UTC · Thursday, 14 August 2025
Duration 47min 14:26 → 15:13 degraded
Peak error rate 6.8% Baseline 0.04% · 170× normal
Checkout p99 9.4s Up from 640 ms steady state
Failed attempts 1,880 Of 12,140 checkouts in window
Orders lost 676 1,204 recovered on retry
Impact

Shoppers on the storefront saw the checkout button hang, then a generic error. Cart contents were preserved and no payment was double-charged. Order capture, fulfillment, and the merchant dashboard were unaffected. Time to detect was 8 minutes; time to mitigate, 35 minutes.

What happened


At 14:02 UTC we began a canary release of checkout-api v4.212.0, which moved tax calculation from a nightly cached rate table to a live quote from our tax vendor. The canary ran clean at 5% of traffic for 17 minutes and was promoted to 100% at 14:19.

The new code called the tax vendor from inside an open Postgres transaction. Under low traffic the extra latency was invisible. At full traffic — roughly 340 checkout requests per second at the Thursday afternoon peak — the average connection hold time rose from 12 ms to 140 ms. The pgbouncer pool has 200 connections; at that hold time, demand exceeded capacity. Requests began queueing for connections, queue wait reached 1.9 s, and the 10-second edge timeout started firing.

We rolled back at 14:52. Latency recovered within three minutes of the rollback reaching 100%, and the error rate returned to baseline at 15:13.

Timeline


  • 14:02UTC
    Canary deploy begins

    checkout-api v4.212.0 goes to 5% of traffic. Release checks pass: error rate flat, p99 at 700 ms.

  • 14:19
    Promoted to 100%

    Automated canary analysis reports no regression. Rollout completes across all 42 pods by 14:24.

  • 14:26
    Degradation starts

    Checkout p99 crosses 2 s. No alert fires — the page threshold is p99 above 4 s sustained for 5 minutes.

  • 14:31
    First customer report

    A support engineer posts two merchant reports of hanging checkouts in #support-escalations.

  • 14:34
    Page fires · acknowledged in 2 min

    checkout_p99_high pages the on-call. S. Okafor acknowledges at 14:36 and opens an investigation.

  • 14:38
    Incident 2041 declared (SEV-3)

    Incident channel and status page draft are created automatically. Two responders join.

  • 14:41
    Pool saturation identified

    pgbouncer shows 200/200 connections active and 1.9 s average client wait. Query time itself is normal at 4 ms p99.

  • 14:47
    Raised to SEV-2

    Error rate crosses 5%. Responders correlate onset with the 14:19 promotion and choose rollback over a forward fix.

  • 14:52
    Rollback begins

    One command reverts to v4.211.4. Pods drain and replace in waves of eight.

  • 15:01
    Rollback at 100% · recovery

    p99 falls to 780 ms within three minutes. Connection wait returns to under 5 ms.

  • 15:13
    Error rate at baseline

    Checkout errors back to 0.04%. Degradation window closes at 47 minutes.

  • 15:40
    Cart recovery sweep

    1,204 shoppers had already retried successfully. Recovery emails go to the remaining 676 preserved carts.

  • 16:20
    Incident resolved

    Sixty minutes of stable metrics. Status page updated and the incident is closed.

Root cause


An external HTTP call was placed inside an open database transaction

v4.212.0 replaced the cached tax rate lookup with a live call to the tax vendor. The call was written inside the existing BEGIN … COMMIT block that reserves the cart, writes the order draft, and decrements inventory. The transaction therefore held a pooled Postgres connection for the full duration of a network round trip.

The vendor is not slow by its own standards: median quote latency was 90 ms and p99 was 1.4 s during the incident, both inside its published SLA. The defect is ours. Holding a connection across any network call converts vendor latency into database pool pressure.

Pool exhaustion is not linear. Below the saturation point, higher hold time is absorbed with no visible effect; above it, wait time climbs steeply. At 340 RPS and 140 ms hold time the theoretical demand was roughly 48 connection-seconds per second against a 200-connection pool — comfortable on paper, but request arrival is bursty, and the p99 tail of the vendor call pushed concurrent demand past 200 for several seconds at a time. Each excursion queued requests that were still queued when the next arrived.

Why the canary did not catch it. At 5% of traffic the service handled about 17 RPS. Connection demand never approached the pool limit, so the regression in hold time produced no change in any signal the canary analysis watches. The canary measured the symptom we expected, not the resource we consumed.

Contributing factors


  • 01
    The canary window was short and the traffic unrepresentative

    Seventeen minutes at 5% is enough to catch an error-rate regression and nothing else. Saturation failures need a traffic share near the knee of the curve to appear at all.

  • 02
    We had no metric for connection hold time

    Dashboards showed query duration, which stayed flat at 4 ms p99 throughout. The quantity that actually changed — how long a request held a connection — was not measured anywhere.

  • 03
    The alert threshold sat above the pain threshold

    Checkout became unusable for many shoppers at 2 s p99. The page fires at 4 s for 5 minutes. Customers reported the problem three minutes before our monitoring did.

  • 04
    The vendor client had no timeout budget

    The SDK default timeout is 10 s — longer than the entire checkout request budget. A slow quote could hold a connection for ten seconds with no circuit breaker to stop it.

  • 05
    The load test used a 5 ms mock of the vendor

    Staging soak tests passed at 400 RPS against a fake that answered instantly, which removed the exact property that caused the outage.

What went well · what was luck


Went well

  • Rollback was a single command and took nine minutes end to end.
  • Responders chose to roll back rather than debug forward, before understanding the cause.
  • Cart state was durable, so 64% of affected shoppers recovered on their own retry.
  • Incident tooling created the channel, timeline, and status page draft without manual work.

Was luck

  • The release landed on a Thursday afternoon, not during a weekend promotion, when peak traffic is 4× higher.
  • A support engineer happened to be reading merchant reports in real time.
  • The vendor's p99 was 1.4 s. At its SLA ceiling of 3 s the pool would have collapsed within four minutes of full rollout.

Action items


ID Action Owner Priority Due Status
AI-2041-1 Move every external call out of the checkout transaction. Quote tax before BEGIN and pass the result in. P. RamanPayments Platform P0 2025-08-20 In progress
AI-2041-2 Emit connection hold time and pool wait as first-class metrics. Page at 250 ms p99 hold time for 3 minutes. S. OkaforObservability P0 2025-08-22 Done
AI-2041-3 Set a 400 ms timeout and a circuit breaker on the tax vendor client, falling back to the cached rate table. D. IversenPayments Platform P1 2025-08-27 In progress
AI-2041-4 Lower the checkout latency page to p99 above 1.5 s for 3 minutes, and add a shopper-facing checkout success-rate SLO. S. OkaforObservability P1 2025-08-29 Done
AI-2041-5 Change canary policy for checkout-path releases: 5% for 30 minutes, then 25% for 30 minutes, with pool metrics in the analysis. M. BellRelease Engineering P1 2025-09-03 Not started
AI-2041-6 Replace vendor mocks in the staging soak with a latency-injecting fake set to the vendor's real p99 of 1.4 s. P. RamanPayments Platform P2 2025-09-10 Not started
AI-2041-7 Add a pool-saturation section to the checkout runbook, including the pgbouncer queries used at 14:41. D. IversenPayments Platform P2 2025-09-05 Done

Note on blame


The change was reviewed by two engineers and passed every gate we had. Nobody involved could have seen this from the diff, because the failure lives in the interaction between a transaction boundary, a vendor's tail latency, and a pool size set four years ago. The fix is not more care during review. It is a metric that makes connection hold time visible, a canary that runs long enough to saturate something, and a rule that keeps network calls out of transactions.