Incident management is where most IT careers begin, which makes it the most practically urgent practice on this trail. The previous waypoint defined an incident — an unplanned interruption, mission: restore service fast. This one walks the actual lifecycle, because "fast" is achieved through structure, not adrenaline.
The lifecycle, stage by stage
- Log everything. Every incident becomes a ticket — even the thirty-second fix. Not bureaucracy: unlogged incidents are invisible to trend-spotting (fifteen "quick fixes" of the same fault is a problem nobody can see), invisible to workload evidence, and invisible to the colleague who hits the same fault tomorrow.
- Categorise. What kind of thing is this — network, hardware, application, access? Categories route tickets to the right team and, later, feed the trend analysis that problem management lives on.
- Prioritise = impact × urgency. The heart of the process. Impact: how widespread is the damage (one user? a department? every customer?). Urgency: how fast does damage compound (payroll cutoff is today vs the report is due next month). The two combine — usually via a simple matrix — into priority. This is why "the CEO's mouse" doesn't automatically outrank "the warehouse scanners are down": one is low impact with loud urgency, the other stops thirty people working. A written matrix converts that argument into arithmetic.
- Investigate, escalate honestly. Two escalation species: functional (this needs deeper expertise — pass to the specialist team) and hierarchic (this needs management awareness or authority — big customer, breached deadline, resources required). Knowing when to escalate is a skill; escalating too late is the classic beginner error. Passing a ticket you can't crack isn't failure — sitting on it for four hours of pride is.
- Resolve, confirm, close. Service restored, user agrees it's restored (a step skipped at your peril), resolution documented in words the next person can actually use. "Fixed" is not documentation.
Major incidents are the special case with its own machinery: separate procedure, a designated coordinator running the response, dedicated communication to stakeholders (the hardest part is honest, regular updates while engineers work), and afterwards — a review that feeds problem management. If the ransomware kill-chain post on this trail is ever your Tuesday, a major incident process is the vehicle you'll live inside.
Two habits separate excellent incident handlers from adequate ones, and both are free. First: communicate proactively — a user updated every hour tolerates a six-hour fix better than one left silent for two. The ticket's status field is not communication. Second: write for the future reader — your resolution notes are tomorrow's knowledge base. The technician who documents well is quietly building the known-error database the whole organisation runs on — and building a reputation with it.