system design courseObservability
Knowing what your system is doing before a customer tells you.
4 chapters16 lessonsabout 1 hour
Record discrete events to answer 'what happened and when' during incidents.
- Structured Logging3 min
- The Log Pipeline2 min
- Sampling and Cost2 min
- What to Log, and What Never to Log3 min
Numeric time-series data to spot trends, alert on anomalies, and drive capacity planning.
- The Four Golden Signals2 min
- Counters, Gauges, and Histograms3 min
- Percentiles, Not Averages2 min
- Dashboards That Help at 3am2 min
Follow a single request across every service to find where time was spent or failures occurred.
- The Lost Request Problem2 min
- Traces, Spans, and Context Propagation3 min
- Sampling Strategies3 min
- Tracing in Practice3 min
Notify on-call when SLO is at risk, before the customer notices.
- Symptoms, Not Causes3 min
- Burn-Rate Alerts3 min
- Killing Alert Fatigue3 min
- Runbooks3 min