Skip to content

16 — Disaster Recovery

Disaster Recovery (DR) is the process of restoring IT infrastructure and systems after a catastrophic failure — natural disaster, cyber attack, data center outage, or region-wide failure.

Analogy: DR is like having an emergency plan for your home. If there’s a fire, you’ve practiced the escape route, have backups of important documents in a safe deposit box, and know where to regroup. You hope you never need it, but you’re ready if you do.


Without a DR plan:

  • Data loss — no backups = lost forever
  • Extended downtime — days or weeks to recover
  • Panic decisions — no plan means chaotic response
  • Regulatory fines — HIPAA, GDPR, PCI-DSS require DR plans
  • Reputation damage — long outages destroy customer trust

flowchart TB
DR["Disaster Recovery<br/>Strategies"] --> Backup["Backup & Restore<br/>Cheapest, slowest recovery"]
DR --> Pilot["Pilot Light<br/>Core data replicated,<br/>servers off until needed"]
DR --> Warm["Warm Standby<br/>Scaled-down version<br/>always running"]
DR --> Multi["Multi-Region<br/>Active-Active<br/>Full production in 2+ regions"]
Backup --> B_Details["RTO: Hours to Days<br/>RPO: 24 hours<br/>Cost: $"]
Pilot --> P_Details["RTO: 10-60 min<br/>RPO: Minutes<br/>Cost: $$"]
Warm --> W_Details["RTO: Minutes<br/>RPO: Seconds<br/>Cost: $$$"]
Multi --> M_Details["RTO: Seconds<br/>RPO: Near-zero<br/>Cost: $$$$"]
style DR fill:#7c3aed,color:#fff
style Backup fill:#3b82f6,color:#fff
style Pilot fill:#059669,color:#fff
style Warm fill:#f59e0b,color:#fff
style Multi fill:#ef4444,color:#fff

StrategyRTORPOCostComplexity
Backup & RestoreHours to days24 hours$Low
Pilot Light10-60 minMinutes$$Medium
Warm StandbyMinutesSeconds$$$Medium
Multi-Region Active-ActiveSecondsNear-zero$$$$High
Data replication onlyDepends on appSeconds$$Medium

The pilot light strategy keeps core data (database, S3) replicated to a DR region. Compute resources (EC2, Lambda) are provisioned only when needed for failover.

flowchart TB
subgraph Primary["Primary Region (us-east-1)"]
APP_PRIMARY["App Servers (running)"]
DB_PRIMARY["Database Primary"]
end
subgraph DR["DR Region (us-west-2)"]
APP_DR["App Servers (stopped)"]
DB_REPLICA["DB Replica (syncing)"]
end
DB_PRIMARY -->|"Continuous Replication"| DB_REPLICA
APP_PRIMARY -.->|"Failover: Start servers,<br/>promote replica"| APP_DR
DB_REPLICA -.->|"Promote to primary"| APP_DR
Users["Users"] --> Route53["Route53<br/>Active-Passive"]
Route53 --> APP_PRIMARY
Route53 -.->|"Health check fails<br/>→ Route to DR"| APP_DR
style Primary fill:#3b82f6,color:#fff
style DR fill:#7c3aed,color:#fff
style DB_PRIMARY fill:#059669,color:#fff
style DB_REPLICA fill:#f59e0b,color:#fff

sequenceDiagram
participant Ops as Operations Team
participant Monitor as Monitoring
participant DR as DR Scripts
participant Cloud as Cloud Provider
Monitor->>Ops: 🚨 Alert: Primary region down!
Ops->>Ops: 📋 Step 1: Verify disaster (5 min)
Ops->>Ops: 📋 Step 2: Declare disaster (1 min)
Ops->>DR: 📋 Step 3: Execute DR plan
DR->>Cloud: Promote DB replica in DR region
DR->>Cloud: Start app servers in DR region
DR->>Cloud: Update DNS to point to DR region
Cloud-->>Ops: DR environment ready ✅
Ops->>Monitor: 📋 Step 4: Verify DR system (15 min)
Ops->>Ops: 📋 Step 5: Update status page
Note over Ops,Cloud: Recovery complete. RTO: 30 min. RPO: 5 min.

Data TypeBackup MethodFrequencyRetention
DatabaseAutomated snapshotsDaily + transaction logs (continuous)30-90 days
Application configVersion control (Git)Per deploymentForever
User uploadsReplication / versioningContinuousRegional replication
LogsExport to S3/GlacierReal-time1-7 years (compliance)
InfrastructureIaC templates (CloudFormation, Terraform)Per changeVersion history

flowchart TB
Rule["3-2-1 Backup Rule"] --> Copies["3 copies<br/>of your data"]
Copies --> Media["2 different media<br/>types (SSD + Tape / Cloud)"]
Media --> Offsite["1 copy offsite<br/>(different region)"]
Copies --> Copy1["Primary (hot)"]
Copies --> Copy2["Local backup (warm)"]
Copies --> Copy3["Offsite backup (cold)"]
style Rule fill:#7c3aed,color:#fff
style Copies fill:#3b82f6,color:#fff
style Media fill:#059669,color:#fff
style Offsite fill:#ef4444,color:#fff

Test TypeDescriptionFrequency
Tabletop exerciseWalk through DR plan verballyQuarterly
Backup restore testActually restore from backupsMonthly
Partial failoverFail over non-critical componentsQuarterly
Full DR drillComplete failover to DR regionAnnually
Chaos engineeringDeliberately inject failuresOngoing

DecisionProsCons
Backup & RestoreCheap, simpleSlow recovery, potential data loss
Pilot LightGood balance of cost and speedManual steps during failover
Warm StandbyFast recoveryRun cost even when not needed
Multi-RegionInstant availabilityHighest cost, complex data sync

StrategyDescription
Automated DR runbookScript everything — no manual steps during disaster
Infrastructure as CodeReproduce entire environment from templates
Data replicationContinuous replication for databases (async or sync)
DNS-based failoverRoute53 health checks + automatic DNS switching
Immutable infrastructureNo manual server configs — rebuild not repair

  1. What’s the difference between pilot light and warm standby?
  2. Explain RTO and RPO — what’s the difference?
  3. What is the 3-2-1 backup rule?
  4. How would you design a disaster recovery plan for a multi-region web app?
  5. How often should you test your disaster recovery plan?

SystemDR Strategy
AWS3 copies of all S3 objects, Geo-redundant across regions
NetflixMulti-region active-active (Simian Army tests DR)
GitHubMySQL replication + S3 backups, DR drills every quarter
GoogleRedundant data centers globally, minutes to failover

  • Disaster Recovery = plan to restore service after a catastrophic failure
  • RTO (Recovery Time Objective) = how fast you recover; RPO = how much data you lose
  • Four strategies from cheapest/slowest to expensive/fastest: Backup → Pilot Light → Warm Standby → Multi-Region
  • Pilot Light = keep DB running in DR region, start app servers on failover — good balance
  • 3-2-1 Rule: 3 copies, 2 media types, 1 offsite
  • Test your DR plan! An untested plan is just a wish
  • Automate everything in the DR runbook — panic leads to mistakes