Vendor-neutral guide · 9 min read
An IT incident response runbook template
Written by the 247connect Marketing Team
Shape of the topic
In short
When something breaks badly, the worst time to decide who is in charge and what counts as serious is during the incident itself. A runbook fixes that in advance: a severity matrix that classifies the incident quickly, named roles so nobody is guessing whose job it is to speak to customers, and a containment sequence that does not depend on the one person who happens to remember it. This template sets out the sections a usable IT incident runbook needs.
Key takeaways
- A severity matrix should be usable within minutes of an incident starting, not require a committee to agree.
- Roles need to be named in advance: incident commander, technical lead and communications lead are not the same job.
- Containment should be documented as a sequence of options, not a single action, since incidents differ.
- Communications need a template and a cadence agreed before an incident, so updates go out even under pressure.
- A post-incident review that produces no changes is a wasted incident; the point is to fix what let it happen.
The severity matrix
The first job of a runbook is classifying an incident fast, so the right level of response kicks in without debate. A simple matrix, scored on business impact and scope, lets whoever picks up the call make that call themselves within minutes rather than escalating a decision about the decision.
Keep the matrix short. A four-level scale that anyone in the team can apply from memory beats a detailed scoring rubric that only gets used properly the second time round, once everyone has forgotten the first incident's stress.
- Severity 1 (Critical): organisation-wide outage or data breach; incident commander engaged immediately, all-hands response
- Severity 2 (High): single site or key system down; on-call technical lead engaged, updates every 30-60 minutes
- Severity 3 (Medium): degraded service affecting a subset of users; handled by standard on-call rota, updates as needed
- Severity 4 (Low): isolated issue, workaround available; logged and handled through normal ticketing
Roles: who does what
Confusion during an incident is rarely about the technical problem, it is about who is deciding what happens next. Naming three roles in advance solves most of it: an incident commander who owns the overall response and makes the call decisions, a technical lead who directs the actual fix, and a communications lead who keeps stakeholders and, where relevant, customers informed.
These roles should be assigned to specific people or an on-call rota before an incident happens, with a clear statement that the incident commander does not need to be the most senior person available, only the one running the response.
- Incident commander: owns the response, makes escalation and containment decisions, declares the incident closed
- Technical lead: directs diagnosis and remediation, reports status to the commander
- Communications lead: manages internal and external updates on an agreed cadence
- Scribe: keeps a timestamped log of actions and decisions for the post-incident review
Containment and initial response
Containment should be documented as a menu of options relevant to common incident types, such as isolating an affected device from the network, disabling a compromised account, or failing over to a backup system, rather than a single generic instruction. The right choice depends on the incident, but having the options pre-written removes the delay of working them out from scratch under pressure.
Whatever action is taken, the scribe role should be logging it with a timestamp as it happens. This log becomes the backbone of the post-incident review and, for security incidents, may be needed as evidence.
Communications during an incident
A pre-agreed communications template, covering what happened, what is affected, what is being done and when the next update will come, means updates go out promptly even when the team is under pressure and short on spare attention. Set the update cadence in the severity matrix so nobody has to decide it mid-incident.
For incidents involving personal data, the communications plan needs to interface with legal and data protection obligations, including the 72-hour notification window to the ICO for qualifying breaches under UK GDPR.
Post-incident review
Close every Severity 1 or 2 incident with a blameless review: what happened, what was done well, what should change, and who owns each resulting action. The review should draw directly on the scribe's timestamped log rather than reconstructed memory, which is unreliable within days of a stressful event.
The review is wasted if it produces no changes. Track the resulting actions to completion the same way you would track any other project, and revisit the runbook itself if the review shows a gap in the process rather than just the technical fix.
Best-practice checklist
1. Publish the severity matrix
Write a short, memorable scale anyone can apply within minutes of an incident starting.
2. Assign roles in advance
Name or roster the incident commander, technical lead, communications lead and scribe before an incident happens.
3. Document containment options
Write a menu of containment actions for common incident types, not a single generic instruction.
4. Agree communications templates and cadence
Pre-write update templates and set the reporting interval per severity level.
5. Set legal and regulatory triggers
Document when data protection or other regulatory notification is required, and who owns that decision.
6. Log everything as it happens
Assign the scribe role to keep a timestamped record of actions and decisions during the incident.
7. Hold a blameless post-incident review
Review every Severity 1 or 2 incident, assign owners to resulting actions, and track them to completion.
8. Test the runbook
Run a tabletop exercise at least annually so the runbook is exercised before a real incident, not during one.
Common pitfalls
- Building a severity matrix so detailed that nobody can apply it quickly under pressure
- Leaving roles undefined, so the most senior person in the room takes over regardless of who should be running the response
- Skipping the communications plan and leaving stakeholders to hear about an outage informally
- Closing an incident without a post-incident review, especially when it felt resolved quickly
- Writing a runbook once and never testing it with a tabletop exercise before a real incident occurs
What to measure
| Time to classify severity | Track minutes from detection to declared severity level |
|---|---|
| Time to first communication | Should meet the cadence set for the declared severity |
| Post-incident reviews completed | Target 100% of Severity 1 and 2 incidents |
| Review actions closed within target | Track completion rate of resulting action items |
| Tabletop exercises run per year | Target at least one for major incident types |
Select any column heading to sort.
Frequently asked questions
- What should be in an IT incident response runbook?
- A usable runbook needs a short severity matrix, named response roles such as incident commander and technical lead, documented containment options for common incident types, a communications template and cadence, and a post-incident review process.
- Who should be the incident commander during an IT incident?
- The incident commander does not need to be the most senior person available; they need to be someone trained to run the response, make containment and escalation decisions, and coordinate the technical and communications leads.
- How quickly should communications go out during an incident?
- This should be set in advance by severity: a critical incident typically needs updates every 30 to 60 minutes, while a lower-severity issue may only need an update when status changes. Pre-agreeing the cadence avoids the decision being made under pressure.
- What is a post-incident review and why does it matter?
- A post-incident review is a blameless examination, held after an incident is resolved, of what happened, what went well, and what should change. It matters because without it, the same underlying cause is likely to produce another incident.
- Does a data breach need to be reported within a specific time frame?
- Under UK GDPR, qualifying personal data breaches must generally be reported to the ICO within 72 hours of the organisation becoming aware of them, which is why the runbook's communications plan should interface directly with legal and data protection obligations.
Sources
Independent, standards-body and peer-reviewed material. None of these sources is affiliated with 247connect.
- Incident Management
NCSC
UK government guidance on preparing for, responding to and learning from cyber incidents.
- Computer Security Incident Handling Guide (SP 800-61 Rev. 2)
NIST
Federal guidance on structuring incident response phases, roles and post-incident activity.
- Personal data breaches
ICO
Guidance on the 72-hour breach notification requirement under UK GDPR.
Putting it into practice
This guide is deliberately product-neutral. If you want to see how one implementation handles these requirements — attended and unattended access, named operator accounts, AES-256 encryption, audit logs and fixed pricing — the reference pages on this hub document 247connect in detail, and the product itself lives at 247connect.cloud.
More best-practice guides
IT asset register template
An IT asset register template covering the fields it needs and the process that keeps it accurate over time.
Remote support acceptable use policy
A remote support acceptable use policy covering consent, what operators may do in a session, recording, monitoring boundaries and breach handling.
MSP client onboarding checklist
A practical MSP client onboarding checklist covering discovery, documentation, agent rollout, escalation paths, the first 30 days and handover to steady state.