The budget reframes the question
Once you have an objective, say 99.9 percent of requests succeeding over 30 days, you own an error budget. A tenth of a percent may fail, which is about 43 minutes of downtime a month.
Let that budget reframe your alerting. Your question is no longer whether the error rate is above some number right now. It is whether, at this failure rate, you will blow the month's budget, and how soon.
Call that spending speed the burn rate. A rate of one means failing at exactly the budgeted pace. A rate of fourteen means the whole month gone in about two days.
See why a plain threshold cannot tell a blip from a crisis. Set it tight and a two minute deploy hiccup pages somebody for a problem that heals itself. Set it loose and a slow simmer of errors quietly eats your month without ever crossing the line.
Fix both ends at once by alerting on burn. Page when it is fast, ticket when it is slow but persistent.
The recipe
Use the standard recipe, which runs several windows. Page at a burn rate of 14 over the last hour, meaning an incident eating 2 percent of your monthly budget every hour. Page at a rate of 6 over six hours. Open a ticket at a rate of 1 over three days, which is the slow leak worth fixing this week.
Add a short confirmation window to each rule, typically five minutes alongside the hour, and fire only when both are burning. The long window proves the problem is sustained and the short one proves it is still happening, which kills the annoying page arriving after a blip already recovered.
Enjoy the proportionality you just bought. A total outage pages within two minutes. A half-percent simmer becomes a calm Tuesday ticket.
Notice too that your thresholds now come from your objective by arithmetic, so tuning stops being folklore and starts being a spreadsheet.
Worked example
Wei's team at a B2B payments company runs a 99.95 percent monthly SLO, about 21 minutes of budget. Their old alert, error rate over 1 percent for 5 minutes, had two famous failures. In February it paged four times in one night for deploy blips that each recovered within 90 seconds, and in April a steady 0.4 percent failure affecting one bank integration ran for six days without a page, burning 80 percent of the budget before a customer escalation surfaced it. Wei implements the SRE Workbook recipe: page at 14.4x burn over 1 hour with a 5-minute confirmation, page at 6x over 6 hours, ticket at 1x over 3 days. The next bank-integration simmer in July generates a ticket on day one, fixed in business hours with 7 percent of budget spent, and night pages drop to one in the following two months.