Site unavailable, no recovery timeline: how to design a DR strategy that survives the loss of a key data center

On the night of October 8, Yandex Cloud experienced a major outage in the ru-central1-b availability zone due to a power failure, leading to a complete shutdown of a key data center. The incident affected numerous services, including Avito, CIAN, and various banking systems. The following day, additional issues were reported at a data center in Kaluga. This situation has prompted a discussion on Disaster Recovery (DR) strategies. The article analyzes infrastructure failure scenarios and proposes architectural solutions to ensure resilience when a primary site is lost. The author examines a recovery model using an independent data center and calculates RPO (Recovery Point Objective) and RTO (Recovery Time Objective) based on a sample infrastructure of 40 virtual machines. The article aims to help engineers prepare for sudden, prolonged cloud provider outages to ensure business continuity under critical conditions.
This is a summary. Read the full article at the original source:
HabrRelated stories
Ukraine’s drones knock out AI data center belonging to "Russia’s Google"
Ukrainian drone strikes have disabled two of the five data centers operated by Yandex, the Russian technology giant often referred to as "Russia’s Goo…
The police drama 'Boston Blue,' a spin-off of the long-running series 'Blue Bloods,' returns for its second season starting October 9. The new season…
Amazon and others are done keeping data center deals secret. Is it enough to build trust?
Amazon has announced it will stop utilizing non-disclosure agreements (NDAs) when negotiating data center projects with local governments, following a…



