Triage
- Customer opens a ticketTier 1
For help desk agents: how to classify, prioritize and escalate a software support ticket from tier 1 to tier 2, engineering or incident management.
Typical practice. Swap the example targets for the ones in your own service level agreement (SLA).
- Acknowledge and log the ticketTier 1
Send the ticket number and the expected response time. Record the customer, account, product, version and best contact method.
- Is something broken, or is it a request?Tier 1
An incident is something broken or degraded. A request is something new, such as access, a new account or a how-to question.
- Request: go to step 4, Fulfil it or route it to the team that owns it
- Broken: go to step 6, Search known issues and the knowledge base
- Fulfil it or route it to the team that owns itTier 1
- Request completed. Close the ticketTier 1
- Search known issues and the knowledge baseTier 1
- Is there a known fix or workaround?Tier 1
- Apply it and walk the customer through itTier 1
Then go to step 34, Confirm the fix with the customer
- Gather the details the next tier will needTier 1
Steps to reproduce, expected and actual result, when it started, account or tenant ID, browser or app version, exact error text, screenshots or logs.
Set the priority
- Set the impactTier 1
High means the whole service, many customers or a critical business function. Medium means several users or one team. Low means one or a few users.
- Set the urgencyTier 1
High means the service is down and no one can work. Medium means a user cannot do their own work. Low means work continues or a workaround exists.
- What priority does the matrix give?Tier 1
High impact with high urgency is P1. High with medium, or medium with high, is P2. High with low, medium with medium, or low with high is P3. The rest are P4.
Priority sets the clocks. One university help desk targets a P1 response in 1 hour and resolution in 8 hours. Use your own SLA.
- P1: go to step 13, Declare a major incident
- P2: go to step 15, Escalate to tier 2 right away
- P3 or P4: go to step 16, Keep working it in tier 1
- Declare a major incidentIncident Manager
Page the on-call incident manager. Open a bridge call, post a status page notice, and update customers on a fixed schedule, for example every 30 minutes.
- Restore service first, then find the root causeEngineering
Then go to step 34, Confirm the fix with the customer
- Escalate to tier 2 right awayTier 2
Then go to step 19, Reproduce the issue and check recent changes
- Keep working it in tier 1Tier 1
- Solved within your tier 1 time limit?Tier 1
Set a short hands-on limit in your SLA. Some desks use 15 to 30 minutes. Use your own limit.
- Escalate to tier 2 with the gathered detailsTier 2
Tier 2 and engineering
- Reproduce the issue and check recent changesTier 2
Look at recent deploys, configuration changes, provider outages, and other tickets with the same symptoms. Link duplicates to one parent ticket.
- Need more information from the customer?Tier 2
- Yes: go to step 21, Ask for everything you need in one message
- No: go to step 24, Can tier 2 fix it?
- Ask for everything you need in one messageTier 2
Many desks pause the SLA clock while waiting on the customer. Send 2 or 3 reminders over several days.
- Did the customer reply?Tier 2
- Close as no response, with a way to reopenTier 2
- Can tier 2 fix it?Tier 2
- Yes: go to step 25, Apply the fix
- No: go to step 26, Is it a product defect?
- Apply the fixTier 2
Then go to step 30, Update the customer at each change and on schedule
- Is it a product defect?Tier 2
- File a bug with engineeringEngineering
Include reproduction steps, impact, priority, affected customers and the ticket link. Engineering confirms the priority or changes it.
- Ship a fix or a workaroundEngineering
Then go to step 30, Update the customer at each change and on schedule
- Hand off to the owning team with a warm handoverTier 2
For example billing, security, or a third-party vendor. Tell the customer who owns it now.
- Update the customer at each change and on scheduleTier 2
Say what you know, what you are doing, and when the next update will come. Typical intervals are every few hours for P2 and daily for P3 and P4.
- Is the resolution target at risk?Tier 2
- Alert the team lead and raise the work's priorityTier 2
Then go to step 34, Confirm the fix with the customer
- Keep working to the targetTier 2
Close
- Confirm the fix with the customerTier 1
- Does the customer confirm it is fixed?Tier 1
- Reopen and send it back to tier 2Tier 1
Then go to step 19, Reproduce the issue and check recent changes
- Close the ticket with the cause and the fix recordedTier 1
- Update or write the knowledge base articleTier 1
Write the symptom in the customer's words, the cause and the fix. Mark it public or internal.
- Ticket closed and knowledge base updatedTier 1