Eighteen months ago the ORBIS Cloud on-call rotation was paging around forty times a month across eight people. It is now fifteen, and we detect real incidents faster than we did.

The change was a rule: a page must correspond to a symptom a customer would notice, and it must be actionable by the person receiving it at 3 a.m. Everything else is a ticket.

Applying that rule deleted most of our alerts. 'Disk usage above 80%' is not a customer symptom and is not actionable at 3 a.m.; it is a ticket, or better, an autoscaling policy. 'Delivery pipeline p99 latency exceeds the SLO error budget burn rate' is both.

We rebuilt alerting around service level objectives and multi-window burn rates. A fast burn pages. A slow burn opens a ticket. The error budget is published and when it is exhausted, feature work stops until reliability work restores it — which has happened twice and was honoured both times, which is the only reason anyone believes the policy.

The unexpected benefit: with fifteen pages a month, people read them. At forty, they had learned not to.