Four signals
Four signals cover service monitoring, and the framing has held up because it starts from user pain instead of machine internals.
Latency is how long your requests take, and you track successful and failed ones separately, because a fast stream of errors will happily drag your average down while everything burns.
Traffic is demand, requests a second for an API or streams for a video service. Errors are the rate of requests failing outright or coming back wrong. Saturation is how full your most constrained resource is: pool usage, queue depth, disk filling.
See why those four earn their place. They cover both directions of causality.
Latency and errors measure the pain your users feel right now. Traffic explains why, since most incidents are load-shaped. Saturation predicts pain, because a queue at 92 percent and climbing is an outage with a countdown attached.
Watch an on-call engineer classify an incident in one glance with all four on screen. Errors up with traffic flat points at a bad deploy or a dying dependency. Errors up with traffic tripled points at load. Everything fine except saturation climbing means you have hours to act, not minutes.
The sibling acronyms
Recognise two sibling acronyms when they turn up in interviews and vendor documentation. One trims the golden signals for request-driven services, easy to apply uniformly across a fleet. The other fits hardware and resources like disks and thread pools.
There is no ideological difference between them, only emphasis. One watches the work arriving and the other watches the thing doing the work.
Instrument all four on every new service before you write a single custom metric. Product-specific numbers matter, and signups a minute will never tell you your database connection pool is exhausted.
Worked example
Ines gets paged at 21:40: checkout error rate at 4 percent and rising. Her service dashboard leads with the four signals, and the shape tells the story in about a minute. Traffic is flat at 3,000 requests per second, so this is not a load event. Latency p99 on calls to the inventory service has jumped from 80 ms to 2 seconds, errors are all timeouts on that same downstream, and saturation shows her service's outbound connection pool pinned at 100 percent. Diagnosis: inventory is slow, her pool is exhausted waiting on it, and requests are timing out in the queue. She flips the feature flag that serves cached stock levels, errors drop to 0.2 percent within 4 minutes, and she escalates to the inventory team with the exact timestamp their latency inflected.