Practice2020

The room where the internet stays up

In a NOC, every incident is a race between two systems: the one that is failing, and the one of people, escalations and handovers trying to understand it. I led the second one.

Team Lead, Network Operations Center — Global-scale cloud provider

Context

In 2020 the world's traffic moved indoors and stayed there. The cloud provider I had joined as an engineer the year before was running infrastructure that suddenly carried everything — work, school, family. I was leading the Network Operations Center team through it.

The problem

A NOC fails in two ways: technically, when an incident outpaces understanding, and organizationally, when the right information exists in the room but never reaches the right person. The second failure is more common and more damaging — and it is a design problem, not a staffing one.

My role

Owning incidents end to end: escalation paths, handovers between shifts, runbooks, and the humans at 3 a.m. Building the operating rhythm — what gets escalated, when, to whom, with what information — and coaching the team through the hardest educational year most of us had worked.

Constraints

Scale meant no single person could hold the system in their head. On-call rotations meant context had to survive handovers intact. Public stakes meant pressure was constant. And the constraint I cared about most: the youngest engineer on shift at 4 a.m. still had to make good decisions with imperfect information.

Discovery

Pattern-reading across incidents taught the durable lessons: which alerts actually predicted trouble versus noise; which escalation paths worked and which just moved anxiety; where runbooks were written for the author instead of the reader. Postmortems were the curriculum — not for assigning cause, but for finding where information had stopped flowing.

The decision

Invest in the information system, not just the technical one: handovers with a fixed structure (what we know, what we've ruled out, what we're watching, who owns next steps), runbooks written to be executed by a tired stranger, and explicit norms that saying 'we don't know yet' is an acceptable status. Ambiguity stated clearly beats false confidence every time.

Trade-offs

Structure costs speed in the easy moments to buy correctness in the hard ones. A fixed handover format feels bureaucratic at 15:00 on a quiet Tuesday and priceless at 03:00 during a multi-region event. Choosing clarity over heroics also means accepting that the brilliant-individual-improvisation path is deliberately closed.

Execution

Working with engineering teams on alert quality, with the shift leads on handover discipline, and with every incident review on feeding the lessons back into the runbooks. Coordination with engineering leadership on what the NOC was seeing before it became their postmortem. The job was equal parts protocol and trust.

Outcome

A team that held through the most demanding year in the platform's traffic history, escalation paths people actually used, and an incident language precise enough that engineering could act on our reports without re-deriving them. The reflection I kept from that year: incident response is applied epistemology.

What I learned

Most 'technical' problems at scale are interface problems between humans. That insight is why I later moved toward product: a product is just an interface between an organization and its users, and the discipline of designing one well is the same.