Is it an incident?
- Issue reported or alert firesService Desk
For service desks and on-call staff: decide if it's an incident, set SEV1 to SEV4 from impact and urgency, then escalate, update and review.
The levels, targets and timings below are examples. Replace them with the ones in your own incident policy.
- Is a service down or working worse than normal?Service Desk
An incident is an unplanned interruption or drop in quality of a service. A request for something new, such as access or a laptop, is a service request.
- Yes: go to step 3, Log it as an incident
- Not sure: go to step 4, Log it as an incident until you know more
- No: go to step 5, Handle it as a service request
- Log it as an incidentService Desk
Record what's affected, when it started, who reported it, how many people are hit, and any error messages.
Then go to step 6, Could it be a security incident?
- Log it as an incident until you know moreService Desk
If you're unsure whether to start the incident process, start it. Closing a false alarm costs less than a slow start.
Then go to step 6, Could it be a security incident?
- Handle it as a service requestService Desk
- Could it be a security incident?Service Desk
Signs: malware or ransomware, unauthorized access, data exposed or stolen, a phishing email that worked, or strange account activity.
- Yes: go to step 7, Page the security on-call team now
- No: go to step 8, Service incident only
- Page the security on-call team nowService Desk
Follow the security incident plan alongside this one. Don't wipe, reboot or power off affected systems unless security tells you to, so evidence survives.
Note the likely type, such as data breach, ransomware, account takeover or denial of service.
Then go to step 9, Assess impact
- Service incident onlyService Desk
Set the severity
- Assess impactService Desk
Impact is how much of the business is affected: number of users, how critical the service is, and any effect on safety, customers or personal data.
- Assess urgencyService Desk
Urgency is how soon the harm grows. A workaround lowers it. A deadline, such as payroll day or a peak sales hour, raises it.
- Which severity fits best?Service Desk
Example matrix: wide impact with high urgency is SEV1. Narrow impact with low urgency is SEV4. Everything else sits between.
If you're torn between two levels, pick the higher one. You can lower it later.
- SEV1: go to step 12, Critical. Core service down or data exposed
- SEV2: go to step 13, High. Important service badly degraded
- SEV3: go to step 14, Medium. Limited impact or a workaround exists
- SEV4: go to step 18, Low. One user or a minor fault
- Critical. Core service down or data exposedService Desk
Examples: a customer-facing service down for everyone, confirmed data exposure, or a safety risk. Treat as a major incident.
Then go to step 20, Declare a major incident
- High. Important service badly degradedService Desk
Examples: many users affected, or a whole site or department down with no workaround. Treat as a major incident.
Then go to step 20, Declare a major incident
- Medium. Limited impact or a workaround existsService Desk
Example target: acknowledge within 1 business day.
- Assign to the owning teamOn call
Page the on-call engineer if the team's runbook says to. Update the ticket and the reporter at least once each business day.
- Is the impact or urgency growing?On call
- Fix, confirm with the reporter and close the ticketOn call
- Low. One user or a minor faultService Desk
Example target: acknowledge within 2 business days. Escalate if more reports of the same fault arrive.
- Resolve in the normal queue and confirm with the userService Desk
Major incident (SEV1 and SEV2)
- Declare a major incidentOn call
Declare when an event meets your written criteria. A major incident needs a coordinated response across teams.
- Page an incident commanderOn call
Use your paging tool. Don't wait until you know the cause. Record the time of the page in the ticket.
- Incident commander takes chargeIncident Commander
The incident commander coordinates and makes the calls, while others do the fixing. Name a scribe to keep the timeline and a lead for communications.
- Open a bridge call and an incident channelIncident Commander
Put the severity, the incident commander's name and the next update time at the top of the channel.
- Send the first updateIncident Commander
Say who's affected, what's known, what's being done and when the next update comes. Use the status page too if customers are affected.
- Send updates on a fixed scheduleIncident Commander
Example cadence: SEV1 every 30 minutes, SEV2 every 60 minutes. Send one on time even when nothing has changed.
- Restore service first, find the root cause laterIncident Commander
Roll back a recent change, fail over, restart or throttle. A workaround that gets users going counts.
Escalate to senior leaders and suppliers if a SEV1 isn't stable by your limit, for example 2 hours after diagnosis starts.
- Is service restored?Incident Commander
- Verify recovery with users and monitoringIncident Commander
Ask affected users to confirm, and watch dashboards for a set time, for example 30 minutes, before you close.
- Send the final update and close the incidentIncident Commander
Give the end time, a short impact summary and whether a review will follow.
Learn from it
- Schedule a blameless post-incident reviewIncident Commander
Example: within 3 calendar days for SEV1 and 5 business days for SEV2. The incident commander names an owner.
Cover the timeline, root cause, customer impact, what went well and action items with owners. Focus on causes, not blame.
- Feed the lessons into runbooks, monitoring and trainingIncident Commander
Lessons learned should improve how you prevent, detect and respond to the next incident.
- Review done. Actions tracked to completionIncident Commander