IT incident severity and escalation: SEV1 to SEV4 flowchart

For service desks and on-call staff: decide if it's an incident, set SEV1 to SEV4 from impact and urgency, then escalate, update and review.

IT incident severity and escalation: SEV1 to SEV4 flowchartService DeskOn callIncident CommanderIS IT AN INCIDENT?SET THE SEVERITYMAJOR INCIDENT (SEV1 AND SEV2)LEARN FROM ITYesNot sureNoYesNoSEV1SEV2SEV3NoSEV4YesYesNoIssue reported or alert firesFor service desks and on-call staff:decide if it's an incident, set SEV1 toSEV4 from impact and urgency, thenescalate, update and review.The levels, targets and timings beloware examples. Replace them with theones in your own incident policy.Is a service down or workingworse than normal?An incident is an unplannedinterruption or drop in quality of aservice. A request for somethingnew, such as access or a laptop, is aservice request.Log it as an incidentRecord what's affected, when itstarted, who reported it, how manypeople are hit, and any errormessages.Log it as an incident until youknow moreIf you're unsure whether to start theincident process, start it. Closing afalse alarm costs less than a slowstart.Handle it as a servicerequestCould it be a securityincident?Signs: malware or ransomware,unauthorized access, data exposedor stolen, a phishing email thatworked, or strange account activity.Page the security on-callteam nowFollow the security incident planalongside this one. Don't wipe, rebootor power off affected systemsunless security tells you to, soevidence survives.Note the likely type, such as databreach, ransomware, accounttakeover or denial of service.Service incident onlyAssess impactImpact is how much of the businessis affected: number of users, howcritical the service is, and any effecton safety, customers or personaldata.Assess urgencyUrgency is how soon the harmgrows. A workaround lowers it. Adeadline, such as payroll day or apeak sales hour, raises it.Which severity fits best?Example matrix: wide impact withhigh urgency is SEV1. Narrow impactwith low urgency is SEV4. Everythingelse sits between.If you're torn between two levels,pick the higher one. You can lower itlater.Critical. Core service downor data exposedExamples: a customer-facing servicedown for everyone, confirmed dataexposure, or a safety risk. Treat as amajor incident.High. Important servicebadly degradedExamples: many users affected, or awhole site or department down withno workaround. Treat as a majorincident.Medium. Limited impact or aworkaround existsExample target: acknowledge within 1business day.Assign to the owning teamPage the on-call engineer if theteam's runbook says to. Update theticket and the reporter at least onceeach business day.Is the impact orurgency growing?Fix, confirm with thereporter and close the ticketLow. One user or a minorfaultExample target: acknowledge within2 business days. Escalate if morereports of the same fault arrive.Resolve in the normal queueand confirm with the userDeclare a major incidentDeclare when an event meets yourwritten criteria. A major incidentneeds a coordinated responseacross teams.Page an incident commanderUse your paging tool. Don't wait untilyou know the cause. Record the timeof the page in the ticket.Incident commander takeschargeThe incident commander coordinatesand makes the calls, while others dothe fixing. Name a scribe to keep thetimeline and a lead forcommunications.Open a bridge call and anincident channelPut the severity, the incidentcommander's name and the nextupdate time at the top of thechannel.Send the first updateSay who's affected, what's known,what's being done and when the nextupdate comes. Use the status pagetoo if customers are affected.Send updates on a fixedscheduleExample cadence: SEV1 every 30minutes, SEV2 every 60 minutes.Send one on time even when nothinghas changed.Restore service first, find theroot cause laterRoll back a recent change, fail over,restart or throttle. A workaround thatgets users going counts.Escalate to senior leaders andsuppliers if a SEV1 isn't stable byyour limit, for example 2 hours afterdiagnosis starts.Is service restored?Verify recovery with usersand monitoringAsk affected users to confirm, andwatch dashboards for a set time, forexample 30 minutes, before youclose.Send the final update andclose the incidentGive the end time, a short impactsummary and whether a review willfollow.Schedule a blamelesspost-incident reviewExample: within 3 calendar days forSEV1 and 5 business days for SEV2.The incident commander names anowner.Cover the timeline, root cause,customer impact, what went well andaction items with owners. Focus oncauses, not blame.Feed the lessons intorunbooks, monitoring andtrainingLessons learned should improve howyou prevent, detect and respond tothe next incident.Review done. Actionstracked to completion

Is it an incident?

  1. Issue reported or alert firesService Desk

    For service desks and on-call staff: decide if it's an incident, set SEV1 to SEV4 from impact and urgency, then escalate, update and review.

    The levels, targets and timings below are examples. Replace them with the ones in your own incident policy.

  2. Is a service down or working worse than normal?Service Desk

    An incident is an unplanned interruption or drop in quality of a service. A request for something new, such as access or a laptop, is a service request.

  3. Log it as an incidentService Desk

    Record what's affected, when it started, who reported it, how many people are hit, and any error messages.

    Then go to step 6, Could it be a security incident?

  4. Log it as an incident until you know moreService Desk

    If you're unsure whether to start the incident process, start it. Closing a false alarm costs less than a slow start.

    Then go to step 6, Could it be a security incident?

  5. Handle it as a service requestService Desk
  6. Could it be a security incident?Service Desk

    Signs: malware or ransomware, unauthorized access, data exposed or stolen, a phishing email that worked, or strange account activity.

  7. Page the security on-call team nowService Desk

    Follow the security incident plan alongside this one. Don't wipe, reboot or power off affected systems unless security tells you to, so evidence survives.

    Note the likely type, such as data breach, ransomware, account takeover or denial of service.

    Then go to step 9, Assess impact

  8. Service incident onlyService Desk

Set the severity

  1. Assess impactService Desk

    Impact is how much of the business is affected: number of users, how critical the service is, and any effect on safety, customers or personal data.

  2. Assess urgencyService Desk

    Urgency is how soon the harm grows. A workaround lowers it. A deadline, such as payroll day or a peak sales hour, raises it.

  3. Which severity fits best?Service Desk

    Example matrix: wide impact with high urgency is SEV1. Narrow impact with low urgency is SEV4. Everything else sits between.

    If you're torn between two levels, pick the higher one. You can lower it later.

  4. Critical. Core service down or data exposedService Desk

    Examples: a customer-facing service down for everyone, confirmed data exposure, or a safety risk. Treat as a major incident.

    Then go to step 20, Declare a major incident

  5. High. Important service badly degradedService Desk

    Examples: many users affected, or a whole site or department down with no workaround. Treat as a major incident.

    Then go to step 20, Declare a major incident

  6. Medium. Limited impact or a workaround existsService Desk

    Example target: acknowledge within 1 business day.

  7. Assign to the owning teamOn call

    Page the on-call engineer if the team's runbook says to. Update the ticket and the reporter at least once each business day.

  8. Is the impact or urgency growing?On call
  9. Fix, confirm with the reporter and close the ticketOn call
  10. Low. One user or a minor faultService Desk

    Example target: acknowledge within 2 business days. Escalate if more reports of the same fault arrive.

  11. Resolve in the normal queue and confirm with the userService Desk

Major incident (SEV1 and SEV2)

  1. Declare a major incidentOn call

    Declare when an event meets your written criteria. A major incident needs a coordinated response across teams.

  2. Page an incident commanderOn call

    Use your paging tool. Don't wait until you know the cause. Record the time of the page in the ticket.

  3. Incident commander takes chargeIncident Commander

    The incident commander coordinates and makes the calls, while others do the fixing. Name a scribe to keep the timeline and a lead for communications.

  4. Open a bridge call and an incident channelIncident Commander

    Put the severity, the incident commander's name and the next update time at the top of the channel.

  5. Send the first updateIncident Commander

    Say who's affected, what's known, what's being done and when the next update comes. Use the status page too if customers are affected.

  6. Send updates on a fixed scheduleIncident Commander

    Example cadence: SEV1 every 30 minutes, SEV2 every 60 minutes. Send one on time even when nothing has changed.

  7. Restore service first, find the root cause laterIncident Commander

    Roll back a recent change, fail over, restart or throttle. A workaround that gets users going counts.

    Escalate to senior leaders and suppliers if a SEV1 isn't stable by your limit, for example 2 hours after diagnosis starts.

  8. Is service restored?Incident Commander
  9. Verify recovery with users and monitoringIncident Commander

    Ask affected users to confirm, and watch dashboards for a set time, for example 30 minutes, before you close.

  10. Send the final update and close the incidentIncident Commander

    Give the end time, a short impact summary and whether a review will follow.

Learn from it

  1. Schedule a blameless post-incident reviewIncident Commander

    Example: within 3 calendar days for SEV1 and 5 business days for SEV2. The incident commander names an owner.

    Cover the timeline, root cause, customer impact, what went well and action items with owners. Focus on causes, not blame.

  2. Feed the lessons into runbooks, monitoring and trainingIncident Commander

    Lessons learned should improve how you prevent, detect and respond to the next incident.

  3. Review done. Actions tracked to completionIncident Commander

Outcomes

Handle it as a service request

You get here from step 2, Is a service down or working worse than normal? (No).

Fix, confirm with the reporter and close the ticket

You get here from step 16, Is the impact or urgency growing? (No).

Resolve in the normal queue and confirm with the user

You get here from step 18, Low. One user or a minor fault.

Review done. Actions tracked to completion

You get here from step 31, Feed the lessons into runbooks, monitoring and training.