Vendor-neutral guide · 9 min read

An IT incident response runbook template

Written by the 247connect Marketing Team

Shape of the topic

A written policy document feeding a numbered checklist, with a review loop returning to the document.A written policy document feeding a numbered checklist, with a review loop returning to the document.
Write it down, work the list, review it: a policy is only useful once it becomes a repeatable checklist.

In short

When something breaks badly, the worst time to decide who is in charge and what counts as serious is during the incident itself. A runbook fixes that in advance: a severity matrix that classifies the incident quickly, named roles so nobody is guessing whose job it is to speak to customers, and a containment sequence that does not depend on the one person who happens to remember it. This template sets out the sections a usable IT incident runbook needs.

Key takeaways

  • A severity matrix should be usable within minutes of an incident starting, not require a committee to agree.
  • Roles need to be named in advance: incident commander, technical lead and communications lead are not the same job.
  • Containment should be documented as a sequence of options, not a single action, since incidents differ.
  • Communications need a template and a cadence agreed before an incident, so updates go out even under pressure.
  • A post-incident review that produces no changes is a wasted incident; the point is to fix what let it happen.

The severity matrix

The first job of a runbook is classifying an incident fast, so the right level of response kicks in without debate. A simple matrix, scored on business impact and scope, lets whoever picks up the call make that call themselves within minutes rather than escalating a decision about the decision.

Keep the matrix short. A four-level scale that anyone in the team can apply from memory beats a detailed scoring rubric that only gets used properly the second time round, once everyone has forgotten the first incident's stress.

  • Severity 1 (Critical): organisation-wide outage or data breach; incident commander engaged immediately, all-hands response
  • Severity 2 (High): single site or key system down; on-call technical lead engaged, updates every 30-60 minutes
  • Severity 3 (Medium): degraded service affecting a subset of users; handled by standard on-call rota, updates as needed
  • Severity 4 (Low): isolated issue, workaround available; logged and handled through normal ticketing

Roles: who does what

Confusion during an incident is rarely about the technical problem, it is about who is deciding what happens next. Naming three roles in advance solves most of it: an incident commander who owns the overall response and makes the call decisions, a technical lead who directs the actual fix, and a communications lead who keeps stakeholders and, where relevant, customers informed.

These roles should be assigned to specific people or an on-call rota before an incident happens, with a clear statement that the incident commander does not need to be the most senior person available, only the one running the response.

  • Incident commander: owns the response, makes escalation and containment decisions, declares the incident closed
  • Technical lead: directs diagnosis and remediation, reports status to the commander
  • Communications lead: manages internal and external updates on an agreed cadence
  • Scribe: keeps a timestamped log of actions and decisions for the post-incident review

Containment and initial response

Containment should be documented as a menu of options relevant to common incident types, such as isolating an affected device from the network, disabling a compromised account, or failing over to a backup system, rather than a single generic instruction. The right choice depends on the incident, but having the options pre-written removes the delay of working them out from scratch under pressure.

Whatever action is taken, the scribe role should be logging it with a timestamp as it happens. This log becomes the backbone of the post-incident review and, for security incidents, may be needed as evidence.

Communications during an incident

A pre-agreed communications template, covering what happened, what is affected, what is being done and when the next update will come, means updates go out promptly even when the team is under pressure and short on spare attention. Set the update cadence in the severity matrix so nobody has to decide it mid-incident.

For incidents involving personal data, the communications plan needs to interface with legal and data protection obligations, including the 72-hour notification window to the ICO for qualifying breaches under UK GDPR.

Post-incident review

Close every Severity 1 or 2 incident with a blameless review: what happened, what was done well, what should change, and who owns each resulting action. The review should draw directly on the scribe's timestamped log rather than reconstructed memory, which is unreliable within days of a stressful event.

The review is wasted if it produces no changes. Track the resulting actions to completion the same way you would track any other project, and revisit the runbook itself if the review shows a gap in the process rather than just the technical fix.

Best-practice checklist

  1. 1. Publish the severity matrix

    Write a short, memorable scale anyone can apply within minutes of an incident starting.

  2. 2. Assign roles in advance

    Name or roster the incident commander, technical lead, communications lead and scribe before an incident happens.

  3. 3. Document containment options

    Write a menu of containment actions for common incident types, not a single generic instruction.

  4. 4. Agree communications templates and cadence

    Pre-write update templates and set the reporting interval per severity level.

  5. 5. Set legal and regulatory triggers

    Document when data protection or other regulatory notification is required, and who owns that decision.

  6. 6. Log everything as it happens

    Assign the scribe role to keep a timestamped record of actions and decisions during the incident.

  7. 7. Hold a blameless post-incident review

    Review every Severity 1 or 2 incident, assign owners to resulting actions, and track them to completion.

  8. 8. Test the runbook

    Run a tabletop exercise at least annually so the runbook is exercised before a real incident, not during one.

Common pitfalls

  • Building a severity matrix so detailed that nobody can apply it quickly under pressure
  • Leaving roles undefined, so the most senior person in the room takes over regardless of who should be running the response
  • Skipping the communications plan and leaving stakeholders to hear about an outage informally
  • Closing an incident without a post-incident review, especially when it felt resolved quickly
  • Writing a runbook once and never testing it with a tabletop exercise before a real incident occurs

What to measure

Metrics for IT incident response runbook
Time to classify severityTrack minutes from detection to declared severity level
Time to first communicationShould meet the cadence set for the declared severity
Post-incident reviews completedTarget 100% of Severity 1 and 2 incidents
Review actions closed within targetTrack completion rate of resulting action items
Tabletop exercises run per yearTarget at least one for major incident types

Select any column heading to sort.

Frequently asked questions

What should be in an IT incident response runbook?
A usable runbook needs a short severity matrix, named response roles such as incident commander and technical lead, documented containment options for common incident types, a communications template and cadence, and a post-incident review process.
Who should be the incident commander during an IT incident?
The incident commander does not need to be the most senior person available; they need to be someone trained to run the response, make containment and escalation decisions, and coordinate the technical and communications leads.
How quickly should communications go out during an incident?
This should be set in advance by severity: a critical incident typically needs updates every 30 to 60 minutes, while a lower-severity issue may only need an update when status changes. Pre-agreeing the cadence avoids the decision being made under pressure.
What is a post-incident review and why does it matter?
A post-incident review is a blameless examination, held after an incident is resolved, of what happened, what went well, and what should change. It matters because without it, the same underlying cause is likely to produce another incident.
Does a data breach need to be reported within a specific time frame?
Under UK GDPR, qualifying personal data breaches must generally be reported to the ICO within 72 hours of the organisation becoming aware of them, which is why the runbook's communications plan should interface directly with legal and data protection obligations.

Sources

Independent, standards-body and peer-reviewed material. None of these sources is affiliated with 247connect.

Putting it into practice

This guide is deliberately product-neutral. If you want to see how one implementation handles these requirements — attended and unattended access, named operator accounts, AES-256 encryption, audit logs and fixed pricing — the reference pages on this hub document 247connect in detail, and the product itself lives at 247connect.cloud.

More best-practice guides