The instinct during a serious outage is to pull in everyone who might be able to help, and beyond a certain point that instinct actively slows recovery. A call with fifteen people, all capable engineers, produces cross-talk, duplicated investigation, and a fix that's harder to coordinate — not because any individual is unhelpful, but because coordination overhead grows faster than the value each additional person adds.
The role that fixes this is an incident commander, separate from whoever is actually diagnosing and fixing the problem. Their job is not technical — it's keeping the call focused, deciding who's actively investigating versus who's on standby, and being the single source of truth for status updates so the person fixing the issue isn't also fielding four different people asking for an update. This is the same principle covered in incident response for small teams, applied at the scale where a formal role becomes necessary rather than an informal one.
Communication cadence matters as much during a major incident as during a smaller one, and it's the part that gets dropped first under pressure. A stakeholder update every fifteen or thirty minutes, even when the update is 'still investigating, no change,' prevents the flood of individual status-check messages that otherwise interrupts the people actually working the problem.
The post-incident review for a major incident deserves more rigour than a routine one — not to assign blame, but because a genuinely major incident usually reveals more than one gap at once: a monitoring blind spot, a runbook that didn't cover this scenario, an escalation path that took too long to activate. Capturing all of them, with owners and dates, is what turns an expensive outage into an investment in not having the next one look the same.