Why Our Sydney AWS Postgres Cluster Failed at Midnight

A deep dive into connection pool exhaustion during a peak Sydney surge and the three config tweaks that saved our database.

POST-MORTEM

9/29/20262 min read

At midnight last Tuesday, our main application cluster in ap-southeast-2 stopped accepting new write requests. What started as a subtle latency bump quickly compounded into total connection pool exhaustion across twelve microservices. Here is the exact breakdown of how our connection routing collapsed and how we brought latency back down to sub-fifty milliseconds.

Diagnosing the Connection Spike in Sydney

Our telemetry flagged an immediate spike in idle connections waiting on lock acquisition. Because our backend services were auto-scaling aggressively in response to batch jobs, each new container spun up its own connection pool directly against the primary instance instead of routing through PgBouncer. Within four minutes, the max connections limit was hit hard.

Refactoring Connection Routing Under Pressure

We immediately adjusted our deployment manifests to enforce strict transaction-level pooling through a centralized proxy layer. We also dialed back maximum container instance counts and updated our health check thresholds so failing pods died before hogging open sockets. Replacing default client timeouts with aggressive five-second caps kept runaway queries from choking the queue.

Key Takeaways for Local Engineering Teams

Never assume default framework database pooling will hold up during rapid traffic spikes. Implementing explicit transaction pooling at the infrastructure layer and testing database behavior under artificial latency in local regional zones remains essential for uptime.