Skip to main content
system design course

Observability

Knowing what your system is doing before a customer tells you.

4 chapters16 lessonsabout 1 hour
0 of 16 lessons readStart the course
  1. 01

    Logging

    0/4

    Record discrete events to answer 'what happened and when' during incidents.

    1. Structured Logging3 min
    2. The Log Pipeline2 min
    3. Sampling and Cost2 min
    4. What to Log, and What Never to Log3 min
  2. Numeric time-series data to spot trends, alert on anomalies, and drive capacity planning.

    1. The Four Golden Signals2 min
    2. Counters, Gauges, and Histograms3 min
    3. Percentiles, Not Averages2 min
    4. Dashboards That Help at 3am2 min
  3. Follow a single request across every service to find where time was spent or failures occurred.

    1. The Lost Request Problem2 min
    2. Traces, Spans, and Context Propagation3 min
    3. Sampling Strategies3 min
    4. Tracing in Practice3 min
  4. Notify on-call when SLO is at risk, before the customer notices.

    1. Symptoms, Not Causes3 min
    2. Burn-Rate Alerts3 min
    3. Killing Alert Fatigue3 min
    4. Runbooks3 min