Support Ticket Categorization Template
Three documents and four sheets that derive your categories from the tickets you actually received, then test every one before you report on it.
Free download · No account needed
Category Register
Havenlea, July 2026 review
Seven candidate categories derived from 13,069 tickets of a shipping and fulfilment queue, then measured. Thresholds set before computing: 30 tickets a month, agreement 0.60, localisation 0.15. Illustrative figures for a fictional company.
| Candidate category | Per month | Agree | Localise | Verdict |
|---|---|---|---|---|
| Label will not print | 590 | 0.79 | 0.14 | split out return labels |
| Order did not sync from the store | 462 | 0.88 | 0.29 | keep |
| Tracking is not updating | 436 | 0.92 | 0.31 | keep |
| Something is not working | 358 | 0.55 | 0.04 | retire |
| Rate at checkout is wrong | 358 | 0.63 | 0.18 | keep |
| Setting up a shipping rule | 315 | 0.69 | 0.42 | keep |
| Customs form incomplete | 95 | 0.94 | 0.20 | keep |
The third largest category failed both gates
| Something is not working, tickets a month | 358 |
| Two coders, agreement corrected for chance | 0.55 |
| Its disposition mix, distance from the queue's | 0.04 |
Nobody could apply it consistently, and knowing a ticket was in it told you nothing you did not know from knowing it was a ticket. It had survived two previous reviews, because it is large and because filing a ticket there is never wrong.
Reasons with no category of their own
| Address validation rejected a good address | 211 | held inside order sync |
| Return label request failed | 192 | held inside label printing |
| Carrier pickup not collected | 161 | held inside tracking |
Five of seven passed every gate. Both failures are instructions rather than deletions, and they are different instructions: one category is retired, one is split.
Every guide on this recommends the same five categories. Every queue running them looks the same: one enormous bucket, one called Other, and the rest mixing what the customer said with what you would have to change. The help desk will not rescue that list. Zendesk's automatic tagging scans each ticket description for words longer than two characters and adds the top three matches against tags already in use, so it reproduces the taxonomy you already had. It cannot surface a category you have no tag for.
This pack derives the candidates from the text of the tickets you received, and lets the data choose how many there are rather than how many fit on a slide. Every candidate then goes through three gates. Enough volume to staff against. Two people independently filing the same ticket in it, against a floor of 0.60, below which methodological work on the agreement statistic treats agreement as inadequate. And a root cause mix that differs from the whole queue's, because a category whose mix is the queue's tells you nothing new.
Havenlea sells shipping and fulfilment software to online retailers. Its inherited five categories put 71.4% of six months of tickets into Technical Issue, which held nine distinct customer problems and five distinct fixes. Seven candidates came out of the derivation and five survived the gates. The third largest category in the queue, 358 tickets a month, failed two gates at once. Two coders agreed on it at 0.55, and its root cause mix sat 0.04 from the whole queue's. It was not a category. It was the absence of one, with a name.
What's in the pack
Category Register
The sheet the taxonomy is decided on, and every later sheet inherits its category list. One row per candidate: the cluster it came from, its highest weighted terms, tickets a month, and its largest root cause. Then the two measurements it had to earn, the agreement figure two independent coders produced and its distance from the queue's own root cause mix, and a verdict. Reasons that got no category of their own are on it too, marked absorbed with the host category and the share of it they represent. A reason the derivation could not separate is a finding rather than an omission.
Volume by Category
The reporting deliverable, on both axes rather than one, with handling hours beside every ticket count and both rankings shown. The two rankings disagree by twenty-two places at the top of the worked example, so a table sorted by count keeps the expensive problem below the fold. The owner sits on the root cause rather than the contact reason, since that is the thing being changed. The metric definitions pack does the same pinning job for every other number in the business. Once a category is ranked by hours, ticket deflection and content gap analysis checks which of them a help center article could actually shrink.
Untagged Queue
The tickets the taxonomy could not place and the ones it placed in two categories at once, with the fit score, the derivation floor it fell short of, and the margin between its top two. A permanent artefact rather than a clean-up list, because it is what the next revision is derived from. Read it by group and not by count: a concentration in one group means a category is missing, and a flat spread means the placement floor is too tight.
Trend
Month by month volume per category across the derivation window, with the held-out month beside the mean and the swing across months on the same row. It exists so a category that only exists because of one month's incident is visible as one, and so a category quietly disappearing from the queue gets retired instead of kept for continuity. The incident that created the spike is easier to trace against the account-level record the customer incident communication template keeps.
Taxonomy Definitions
What each category means, written so somebody holding a ticket can decide. Every definition carries the two figures it had to earn, and a category with no figures next to it is not in force. It also records what was retired and why, in enough detail that the argument for keeping a large useless category only has to be had once.
Tagging Guidance
When to set each field, why you categorise from the first inbound message rather than the thread, and the tie-break rules for the pairs that genuinely overlap, each written as a question about the ticket a person can answer. Also the four separate reasons the tag field is not the category field, which is worth reading before anybody reports off tags. Pairs with the naming convention pack when the fix is the namespace rather than the taxonomy.
Review Procedure
What happens monthly, quarterly and annually: read the unplaced by group, recompute the three gates, and once a year re-derive from scratch with no reference to the taxonomy in force. Every revision has to reconcile volume before and after, because a revision where the total does not match is a revision that lost tickets. Its sibling support ticket export template covers getting the data out completely in the first place.
How it works
- 1
Send the tickets, not the summary
Subjects and bodies, or the first inbound message on each conversation, six months if you have it. River holds the most recent month back before reading anything else, because the taxonomy has to be tested on tickets it was not derived from. Your current category list is useful as the comparison, not as the starting point.
- 2
Derive the candidates and sweep the count
Term weighting, then drop the terms on more than half the tickets, which is where the greetings and your own product name go. Then sweep the number of categories and read the separation. The worked example had ten distinct contact reasons in the queue and seven separated best, and that gap is the useful part rather than a rounding problem.
- 3
Run the three gates per category
Volume, agreement between two independent coders, and distance from the queue's own root cause mix. Per category and not just overall, because an overall figure of 0.77 routinely contains a category sitting at 0.60 by itself. Every failure comes back as an instruction with the tickets it held reconciled somewhere accountable.
- 4
Test it on the month it never saw
Both thresholds derived from the training distribution rather than chosen, then three counts: placed in exactly one category, matched two, placed in none. The double matches get grouped by pair, so the largest pair becomes one tie-break rule rather than a few hundred individually contested tickets.
Frequently asked questions
Why not just use the five standard support categories?
Because the biggest of them is not a category. On the worked example Technical Issue held 71.4% of the queue, containing nine distinct customer problems and five distinct fixes, with handling times spread almost exactly like the whole queue's. Grouping by those five categories accounted for 4.3% of the variation in handling time.
How many categories should a support taxonomy have?
It is a question for your tickets, not for a guide. Sweep the count and read the separation across the range. Expect fewer than the number of problems you can name: three of the worked example's ten shared their vocabulary too closely for anyone reading a ticket to separate, so splitting them would have failed on the queue.
Can't Zendesk's automatic tagging do this for me?
No, and its documentation says why. It matches ticket words against tags already in use, so it reproduces the taxonomy you have and cannot surface one you lack. It also skips tickets agents raise inside the help desk and may not work outside English, so it produces a systematically untagged population.
What makes a category worth keeping?
Three things, measured before it goes into the definitions. Enough volume to staff against. Two people independently filing the same ticket in it. And a root cause mix that differs from the whole queue's, because a category whose mix is the queue's mix tells you nothing new. Five of seven candidates passed all three.
What do I do about the Other bucket?
Run both gates on it and publish the figures rather than arguing. On the worked example it was the third largest category at 358 tickets a month, agreed at 0.55 against a 0.60 floor, and sat 0.04 from the queue's own root cause mix. It was retired, and the residue moved to a visible sheet.
Should the same field hold the customer's problem and the root cause?
No. They are close to independent, at 0.108 normalised mutual information on the worked example, so one column carries one answer and discards the other. Report volume on the pair. The sales and marketing SLA pack makes the same argument about a handoff whose two sides get measured as one.
How do I know the taxonomy will hold up next month?
Apply it to a month it was not derived from and publish three counts: placed in exactly one category, matched two, placed in none. The worked example placed 83.9%, double-matched 11.7% and could not place 4.4%. A category list quoting none of those has never been tested on tickets it had not already seen.