Skip to main content
system design course

Reliability

Designing for the day something fails, because something will.

5 chapters17 lessonsabout 1 hour
0 of 17 lessons readStart the course
  1. No single point of failure. A backup is always ready to take over.

    1. Single Points of Failure2 min
    2. Active-Passive and Active-Active2 min
    3. Health Checks and Split Brain3 min
    4. Failover Drills2 min
  2. Executing the same operation multiple times has the same effect as executing it once.

    1. Why Retries Need Idempotency2 min
    2. Idempotency Keys2 min
    3. Designing Idempotent Operations3 min
  3. Plan for full region failure. RTO and RPO define how fast and how much data you can afford to lose.

    1. RTO and RPO3 min
    2. Backups That Actually Restore3 min
    3. DR Architectures by Budget3 min
  4. SLI measures reality, SLO sets the target, SLA is the promise to customers.

    1. SLI, SLO, SLA: Measure, Target, Promise3 min
    2. Error Budgets2 min
    3. Choosing Good SLOs3 min
  5. Getting a group of machines to agree on one value, so failover cannot produce two primaries.

    1. Split Brain3 min
    2. Quorums2 min
    3. How Raft Elects a Leader3 min
    4. Using It Without Building It3 min