Alerting
Notify on-call when SLO is at risk, before the customer notices.
At 03:14 the pager goes off. Processor usage on one machine crossed 80 percent. The on-call engineer acknowledges it from bed, half asleep, and goes back to sleep, because it is the fourth one this week and none of them meant anything.
At 04:02 it goes off again, and they acknowledge that one the same way. That one was checkout returning errors to every customer.
Nobody was negligent. Forty pages a week, almost all of them noise, trains a person to dismiss pages, and the reflex cannot tell the fortieth false alarm from the real thing.
An alert strategy is a filter: out of everything your telemetry could say, which few conditions are worth waking somebody for. Get it wrong one way and people stop trusting the pager. Get it wrong the other and nobody hears about the outage until morning.
Lessons
4 in this chapter- Symptoms, Not CausesPage on what users feel; graph the causes for the investigation afterward.3 min
- Burn-Rate AlertsAlert on how fast you are spending error budget, so severity sets urgency automatically.3 min
- Killing Alert FatigueEvery page that needs no action teaches the on-call to ignore the one that does.3 min
- RunbooksThe difference between a 10-minute page and a 2-hour one is usually a document.3 min