Incident Command for Small Teams: Adapting ICS to Five Engineers

Incident Command for Small Teams: Adapting ICS to Five Engineers

Reading time1 min
#devops#startups#incident-management#infrastructure#engineering

Incident Command for Small Teams: Adapting ICS to Five Engineers

The Incident Command System (ICS) is the structure American emergency services use to run wildfires, floods and other large incidents. FEMA teaches it as part of the National Incident Management System. The incident process in Google's SRE book is explicitly based on it.

The core idea is simple: separate the person who coordinates from the people doing hands-on work, and give every function a named owner. That works with five engineers as well as with five hundred firefighters, as long as you keep only the parts that fit.

What ICS actually defines

Full ICS has an Incident Commander, a command staff (Public Information Officer, Safety Officer, Liaison Officer) and four general staff sections: Operations, Planning, Logistics, and Finance/Administration. A software team does not need the boxes. It needs the principles behind them:

  • The Incident Commander owns every function that has not been delegated. If nobody is assigned to communication, the IC does it.
  • Modular organization. Activate only the roles the incident needs and add more as it grows.
  • Unity of command. Each person takes direction from one person, so nobody gets conflicting instructions.
  • Explicit transfer of command. Command changes hands with a briefing and an acknowledgement, never by drift.
  • Common terminology. The same words for severity, status and roles every time.

The small-team version

Most software incidents need four roles at most. The SRE book uses almost the same set: incident command, operational work, communication and planning.

Incident Commander (IC). Coordinates, decides, and keeps the overall picture. The IC does not debug. The moment the IC opens a terminal, nobody is watching the whole incident. By default the IC is whoever is on call when the incident is declared. If that person is the one who knows the broken system best, they hand command to someone else and go fix it.

Responders. The people changing things in production. Every change has one owner, and they announce it before doing it: "rolling back my-app to the previous release in the production cluster".

Communications. Posts status updates to the internal channel, the status page and customer-facing colleagues on a fixed cadence. Until someone is assigned, this is the IC's job.

Scribe. Keeps the timeline: what was seen, what was tried, what was decided, and when. It makes the postmortem much easier to write. On a small team the IC can do this by keeping all decisions in the incident channel and treating it as the log.

With two people available, run IC plus one responder. With one person, they hold every role and their first action is to page someone else.

Declaring an incident

A lot of time is lost before anyone admits there is an incident. Write down simple triggers, for example:

  • customer-facing impact,
  • more than one person needs to be involved,
  • the problem is still not understood after a fixed amount of focused debugging (the SRE book suggests one hour).

Declaring is cheap: if it turns out to be minor, cancelling costs one message. Not declaring a real incident costs coordination for its whole duration.

Define three or four severity levels in plain language, based on impact (who is affected and how badly), not on cause. Every incident gets a severity at declaration, and the IC can change it later.

Communication rules

  • One channel per incident with a predictable name, such as #inc-<date>-<short-name>.
  • Decisions and actions go in the channel, not in DMs or side calls.
  • Status updates go out on the agreed cadence even when nothing changed. "No change, still investigating, next update at 14:30 UTC" is a valid update.
  • Handoffs are explicit. The outgoing IC briefs the incoming one, and the incoming IC confirms in the channel that they now hold command.

Creating the channel can be one command, so nobody skips it. This uses the Slack Web API and needs a bot token with the channels:manage scope:

#!/usr/bin/env bash
set -euo pipefail
name="inc-$(date -u +%Y%m%d)-${1:?usage: $0 short-name}"
curl -s -X POST https://slack.com/api/conversations.create \
  -H "Authorization: Bearer ${SLACK_BOT_TOKEN}" \
  -H "Content-Type: application/json; charset=utf-8" \
  -d "{\"name\": \"${name}\"}" | jq -r 'if .ok then .channel.id else .error end'

Slack returns HTTP 200 for most errors, which is why the script checks the ok field instead of the status code.

The incident document

Keep one short living document per incident. The IC owns it. A template that fits on one screen:

# Incident: <short name>
Severity: SEV2            Status: investigating | mitigated | resolved
IC: @name   Responders: @name   Comms: @name   Scribe: @name
Started: <UTC time>       Declared: <UTC time>

## Impact
Who is affected and how. Update as it changes.

## Current hypothesis
One or two sentences.

## Actions in flight
- @owner: action, expected result, when to check back

## Timeline (UTC)
- HH:MM alert fired
- HH:MM incident declared, IC @name

Closing the incident

The IC declares the incident mitigated when impact stops, and resolved when the system is back to a stable state. Before closing, the IC assigns a postmortem owner and a date. Without that step the follow-up work rarely happens.

Practice the process when nothing is broken. A short game day where someone plays IC on a simulated outage shows quickly whether people know the roles, where the template lives and how to open a channel.

Checklist

  • IC is a role in the on-call rotation, and the IC does not debug.
  • Every function has an owner; anything unassigned belongs to the IC.
  • Written triggers for declaring an incident and plain-language severity levels.
  • One channel and one document per incident, created by a script.
  • Status updates on a fixed cadence, handoffs acknowledged in the channel.
  • A postmortem owner and date assigned before the incident is closed.