Technologies
Back
Cloud Computing & Infrastructure

Site unavailable, no recovery timeline: how to design a DR strategy that survives the loss of a key data center

Habr
Advertisement468 × 90
Site unavailable, no recovery timeline: how to design a DR strategy that survives the loss of a key data center

On the night of October 8, Yandex Cloud experienced a major outage in the ru-central1-b availability zone due to a power failure, leading to a complete shutdown of a key data center. The incident affected numerous services, including Avito, CIAN, and various banking systems. The following day, additional issues were reported at a data center in Kaluga. This situation has prompted a discussion on Disaster Recovery (DR) strategies. The article analyzes infrastructure failure scenarios and proposes architectural solutions to ensure resilience when a primary site is lost. The author examines a recovery model using an independent data center and calculates RPO (Recovery Point Objective) and RTO (Recovery Time Objective) based on a sample infrastructure of 40 virtual machines. The article aims to help engineers prepare for sudden, prolonged cloud provider outages to ensure business continuity under critical conditions.

This is a summary. Read the full article at the original source:

Habr
Advertisement468 × 90
Share
Cloud Computing & Infrastructure

Related stories

Advertisement970 × 250