+91 98726 60544 hello@mitstech.co Mon–Sat · 09:00–18:30 IST

On-call without burning out the team

Cloud By Mits Engineering Team 2 min read
On-call without burning out the team

On-call is where reliability practice meets human cost, and the cost is usually paid quietly until someone resigns. The failure is rarely the existence of a rota. It is a rota where the pages are frequent, unactionable, or arrive for systems the person on duty has no ability to fix, and where the hours spent awake at night are treated as free.

The first principle is that every page must be actionable. If the response to an alert is to look at it, confirm nothing is wrong, and go back to sleep, that alert has cost a night's sleep and produced nothing. It should be a dashboard. Audit your alerts quarterly and delete anything that has never once resulted in an action - the deletion improves reliability, because an engineer who trusts the pager responds faster than one who has learned to assume it is noise.

The second is that the person on call must be able to act. That means a runbook for every alert - what it means, what to check, what to do, who to escalate to - written by whoever built the system rather than by whoever is on duty at three in the morning. It also means permissions: an engineer who can diagnose but cannot restart, roll back or scale is not on call, they are a relay to someone who is.

On the shape of the rota, weekly rotations work better than daily because they let people plan around it, and a minimum of four or five people in the rotation is what makes it sustainable - with three, every third week is disrupted and nobody ever fully recovers. Secondary on-call matters as much as primary: the primary needs to know that someone will pick up if they are driving, in a queue, or simply asleep through the first page.

Compensate it explicitly, in money or in time off, and be seen to do so. Unpaid on-call is a tax levied on the most conscientious people in the team, and they notice. Equally important, and more often skipped: if someone was up at two in the morning, they are not expected online at nine. Make that a rule rather than a kindness, because otherwise nobody takes it.

The measure of whether any of this is working is the number of out-of-hours pages per rotation, tracked over time. If it is not falling, the reliability work is not landing, and the on-call burden is being used as a substitute for fixing things. Treat repeat pages as bugs with an owner - the third time a given alert fires at night, the fix goes into the sprint rather than the backlog.

Need help with this? Explore our Cloud Solutions & Migration services. Learn more Back to all news

Keep reading

More on Cloud