Support ticket escalation process: tier 1 to tier 2, engineering and incidents

For help desk agents: classify a ticket, set priority by impact and urgency, and know when to escalate to tier 2, engineering or a major incident.

Support ticket escalation process: tier 1 to tier 2, engineering and incidentsTier 1Incident ManagerEngineeringTier 2TRIAGESET THE PRIORITYTIER 2 AND ENGINEERINGCLOSERequestBrokenYesNoP1P2P3 or P4NoYesNoNoYesNoYesNoYesNoNoYesYesYesCustomer opens a ticketFor help desk agents: how to classify,prioritize and escalate a softwaresupport ticket from tier 1 to tier 2,engineering or incident management.Typical practice. Swap the exampletargets for the ones in your ownservice level agreement (SLA).Acknowledge and log theticketSend the ticket number and theexpected response time. Record thecustomer, account, product, versionand best contact method.Is something broken, or is it arequest?An incident is something broken ordegraded. A request is somethingnew, such as access, a new accountor a how-to question.Fulfil it or route it to the teamthat owns itRequest completed. Closethe ticketSearch known issues andthe knowledge baseIs there a known fixor workaround?Apply it and walk thecustomer through itGather the details the nexttier will needSteps to reproduce, expected andactual result, when it started,account or tenant ID, browser or appversion, exact error text, screenshotsor logs.Set the impactHigh means the whole service, manycustomers or a critical businessfunction. Medium means severalusers or one team. Low means one ora few users.Set the urgencyHigh means the service is down andno one can work. Medium means auser cannot do their own work. Lowmeans work continues or aworkaround exists.What priority does thematrix give?High impact with high urgency is P1.High with medium, or medium withhigh, is P2. High with low, mediumwith medium, or low with high is P3.The rest are P4.Priority sets the clocks. Oneuniversity help desk targets a P1response in 1 hour and resolution in 8hours. Use your own SLA.Declare a major incidentPage the on-call incident manager.Open a bridge call, post a status pagenotice, and update customers on afixed schedule, for example every 30minutes.Restore service first, thenfind the root causeEscalate to tier 2 right awayKeep working it in tier 1Solved within your tier 1 timelimit?Set a short hands-on limit in yourSLA. Some desks use 15 to 30minutes. Use your own limit.Escalate to tier 2 with thegathered detailsReproduce the issue andcheck recent changesLook at recent deploys, configurationchanges, provider outages, and othertickets with the same symptoms.Link duplicates to one parent ticket.Need moreinformation from thecustomer?Ask for everything you needin one messageMany desks pause the SLA clockwhile waiting on the customer. Send2 or 3 reminders over several days.Did the customerreply?Close as no response, with away to reopenCan tier 2 fix it?Apply the fixIs it a productdefect?File a bug with engineeringInclude reproduction steps, impact,priority, affected customers and theticket link. Engineering confirms thepriority or changes it.Ship a fix or a workaroundHand off to the owning teamwith a warm handoverFor example billing, security, or athird-party vendor. Tell the customerwho owns it now.Update the customer at eachchange and on scheduleSay what you know, what you aredoing, and when the next update willcome. Typical intervals are every fewhours for P2 and daily for P3 and P4.Is the resolutiontarget at risk?Alert the team lead and raisethe work's priorityKeep working to the targetConfirm the fix with thecustomerDoes the customerconfirm it is fixed?Reopen and send it back totier 2Close the ticket with thecause and the fix recordedUpdate or write theknowledge base articleWrite the symptom in the customer'swords, the cause and the fix. Mark itpublic or internal.Ticket closed and knowledgebase updated

Triage

  1. Customer opens a ticketTier 1

    For help desk agents: how to classify, prioritize and escalate a software support ticket from tier 1 to tier 2, engineering or incident management.

    Typical practice. Swap the example targets for the ones in your own service level agreement (SLA).

  2. Acknowledge and log the ticketTier 1

    Send the ticket number and the expected response time. Record the customer, account, product, version and best contact method.

  3. Is something broken, or is it a request?Tier 1

    An incident is something broken or degraded. A request is something new, such as access, a new account or a how-to question.

  4. Fulfil it or route it to the team that owns itTier 1
  5. Request completed. Close the ticketTier 1
  6. Search known issues and the knowledge baseTier 1
  7. Is there a known fix or workaround?Tier 1
  8. Apply it and walk the customer through itTier 1

    Then go to step 34, Confirm the fix with the customer

  9. Gather the details the next tier will needTier 1

    Steps to reproduce, expected and actual result, when it started, account or tenant ID, browser or app version, exact error text, screenshots or logs.

Set the priority

  1. Set the impactTier 1

    High means the whole service, many customers or a critical business function. Medium means several users or one team. Low means one or a few users.

  2. Set the urgencyTier 1

    High means the service is down and no one can work. Medium means a user cannot do their own work. Low means work continues or a workaround exists.

  3. What priority does the matrix give?Tier 1

    High impact with high urgency is P1. High with medium, or medium with high, is P2. High with low, medium with medium, or low with high is P3. The rest are P4.

    Priority sets the clocks. One university help desk targets a P1 response in 1 hour and resolution in 8 hours. Use your own SLA.

  4. Declare a major incidentIncident Manager

    Page the on-call incident manager. Open a bridge call, post a status page notice, and update customers on a fixed schedule, for example every 30 minutes.

  5. Restore service first, then find the root causeEngineering

    Then go to step 34, Confirm the fix with the customer

  6. Escalate to tier 2 right awayTier 2

    Then go to step 19, Reproduce the issue and check recent changes

  7. Keep working it in tier 1Tier 1
  8. Solved within your tier 1 time limit?Tier 1

    Set a short hands-on limit in your SLA. Some desks use 15 to 30 minutes. Use your own limit.

  9. Escalate to tier 2 with the gathered detailsTier 2

Tier 2 and engineering

  1. Reproduce the issue and check recent changesTier 2

    Look at recent deploys, configuration changes, provider outages, and other tickets with the same symptoms. Link duplicates to one parent ticket.

  2. Need more information from the customer?Tier 2
  3. Ask for everything you need in one messageTier 2

    Many desks pause the SLA clock while waiting on the customer. Send 2 or 3 reminders over several days.

  4. Did the customer reply?Tier 2
  5. Close as no response, with a way to reopenTier 2
  6. Can tier 2 fix it?Tier 2
  7. Apply the fixTier 2

    Then go to step 30, Update the customer at each change and on schedule

  8. Is it a product defect?Tier 2
  9. File a bug with engineeringEngineering

    Include reproduction steps, impact, priority, affected customers and the ticket link. Engineering confirms the priority or changes it.

  10. Ship a fix or a workaroundEngineering

    Then go to step 30, Update the customer at each change and on schedule

  11. Hand off to the owning team with a warm handoverTier 2

    For example billing, security, or a third-party vendor. Tell the customer who owns it now.

  12. Update the customer at each change and on scheduleTier 2

    Say what you know, what you are doing, and when the next update will come. Typical intervals are every few hours for P2 and daily for P3 and P4.

  13. Is the resolution target at risk?Tier 2
  14. Alert the team lead and raise the work's priorityTier 2

    Then go to step 34, Confirm the fix with the customer

  15. Keep working to the targetTier 2

Close

  1. Confirm the fix with the customerTier 1
  2. Does the customer confirm it is fixed?Tier 1
  3. Reopen and send it back to tier 2Tier 1

    Then go to step 19, Reproduce the issue and check recent changes

  4. Close the ticket with the cause and the fix recordedTier 1
  5. Update or write the knowledge base articleTier 1

    Write the symptom in the customer's words, the cause and the fix. Mark it public or internal.

  6. Ticket closed and knowledge base updatedTier 1

Outcomes

Request completed. Close the ticket

You get here from step 4, Fulfil it or route it to the team that owns it.

Close as no response, with a way to reopen

You get here from step 22, Did the customer reply? (No).

Ticket closed and knowledge base updated

You get here from step 38, Update or write the knowledge base article.