When the cooling died in Proton’s Frankfurt data center, the on-call engineers made a call that almost never has to be made: prioritize saving the hardware, even if that might extend the outage. The company’s incident report, published a day after the August 27 outage by CTO Bart Butler, lays out why. Room temperature climbed from a nominal 21.8 to 51.9 degrees Celsius in under half an hour, and in a market strained by the AI boom, much of the equipment inside, if lost, “would not be possible to replace on short timelines.”

Key facts

  • Trigger: total cooling failure in the main room of Proton’s Frankfurt data center, just after 11 p.m. CEST on August 26.
  • User impact: from around midnight; most services back for most users by 01:30, with less critical systems such as push notifications and payment processing following around 02:00. No email lost, delivery delayed both ways.
  • Root cause: an air filter replacement performed on both redundant air compressors by the data center operator, at night, without prior notice.
  • Losses: some servers suffered what the report calls “heat death”; the lifespan effect on surviving hardware is unknown.

Half an Hour From Nominal to 51.9 Degrees

The failure began just after 11 p.m. on Wednesday, August 26. By about 11:15, temperatures were climbing, reaching 51.9 degrees in under thirty minutes, with some probes reading 60 degrees of air temperature. Equipment began dying one machine at a time. Around midnight the failures crossed the line into a user-facing incident: both the primary and the backup network switch on one critical rack failed, and that rack held several primary database copies.

Proton’s primary databases do not fail over automatically; the company keeps a human in the loop to avoid “split brain” divergence between database copies. Standard procedure would fail them over within the same data center, but with Frankfurt overheating, replicas in the same data center might be next to die. The failover logic, the report concedes, handles a complete data center failure well and a scenario where random servers die one by one much worse.

The Choice, and a Recovery Harder Than Expected

The temperature curve made the decision. In the report’s words, “what used to take 3-4 hours to go critical went critical in 20 minutes,” a consequence of the power density that modern CPUs and GPUs have brought to racks. And because the AI boom has made server hardware scarce, much of the equipment, had it been lost, could not have been replaced quickly. The on-call team spent the critical window powering machines down to protect them and working with the site’s operations team to bring cooling back, accepting a potentially longer outage as the price.

Cooling returned by 00:45, and recovery brought its own lesson: many network cards had reached 105 degrees, far above their 45-degree operating norm, tripping a thermal protection mode that only a cold reset clears. Proton’s own security posture, which restricts access to out-of-band controllers, meant waking additional staff to assist the recovery. Most services returned by 01:30, with push notifications and payment processing following around 02:00, and the database team then spent the night and the following day rebuilding full redundancy across Frankfurt and Zurich.

A Filter Change at Night, on Both Compressors

The root cause, established the next day, was almost mundane: the data center operator replaced air filters on both redundant air compressors powering the cooling system, in the middle of the night, without prior notice. The operator also failed to communicate the cooling failure once it happened, which, in the report’s words, “dramatically cut down the time we had to respond.” Proton says it is working with the operator to prevent a repeat, and describes the series of events that led to the incident as “highly improbable.”

The report is equally direct about Proton’s own side: its database infrastructure has a known limitation that can make this type of outage slower to recover from, with resilience work underway and planned for completion by the end of the year. New data center capacity is being commissioned and is expected to come online in the coming weeks, further reducing single-site dependency. The incident, as Butler puts it, arrived before those improvements were fully in place.

Two Cooling Failures in One August, and the Lessons That Travel

For hosting and infrastructure operators, the report reads as a checklist of assumptions worth revisiting. Redundant switches that share a rack share a fate. Manual failover policies priced for a calm hour cost more in a hot one. Density has quietly rewritten thermal response windows, and hardware scarcity now belongs in incident priorities, not just procurement plans. It is also the second publicly documented cooling-caused outage to hit a major provider in August, by our count: two weeks earlier, a cooling system failure at a Phoenix data center took Namecheap’s website and a number of its services offline, as the company’s own account records. Cooling has had a very visible month, and Proton’s report shows how rising rack density can cut the time to respond from hours to minutes.

About the Data

All timeline details, temperatures, quotes and remediation plans come from Proton’s incident report of August 28, 2026, written by CTO Bart Butler and read in full. Proton’s status page documents the incident as it was communicated overnight. The August 13 Phoenix cooling failure is documented in Namecheap’s own outage account. The closing lessons are our reading of the report’s facts.