· · 1 min · infrastructure · by the wire desk

The chillers failed in Northern Virginia

May 7 and 8, multiple cooling units down at once, one availability zone thermally offline. The dependency map did the rest.

On May 7 and 8, multiple chiller units failed simultaneously in an AWS Northern Virginia data center, triggering thermal-safety shutdowns across EC2 instances and EBS volumes in availability zone use1-az4, per incident tracking. Engineers restored cooling before hardware could be safely re-energized, which is the correct order and the slow one.

Thermal events are the least software-shaped failures in the catalog. There is no rollback for heat. Recovery time is bounded by physics and electrician scheduling, and the incident is a reminder that beneath every availability abstraction sits a building, and buildings have plumbing.

The zone-level containment mostly worked as designed, which makes this a test the industry keeps grading itself on generously: workloads genuinely spread across zones shrugged; workloads that had quietly concentrated in one zone, or depended on something that had, rediscovered the difference between the architecture diagram and the deployment. us-east-1's gravitational pull makes that rediscovery a recurring seminar.

The builder's read: multi-AZ is a claim you verify, not a checkbox you set. Kill a zone in staging this quarter and watch what actually happens. The chillers ran the test in production for whoever had not.

tags: #aws #outage #us-east-1 #datacenters