River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Support Ticket Categorization Template

Three documents and four sheets that derive your categories from the tickets you actually received, then test every one before you report on it.

Free download  ·  No account needed

Category Register

Havenlea, July 2026 review

Seven candidate categories derived from 13,069 tickets of a shipping and fulfilment queue, then measured. Thresholds set before computing: 30 tickets a month, agreement 0.60, localisation 0.15. Illustrative figures for a fictional company.

Candidate categoryPer monthAgreeLocaliseVerdict
Label will not print5900.790.14split out return labels
Order did not sync from the store4620.880.29keep
Tracking is not updating4360.920.31keep
Something is not working3580.550.04retire
Rate at checkout is wrong3580.630.18keep
Setting up a shipping rule3150.690.42keep
Customs form incomplete950.940.20keep

The third largest category failed both gates

Something is not working, tickets a month358
Two coders, agreement corrected for chance0.55
Its disposition mix, distance from the queue's0.04

Nobody could apply it consistently, and knowing a ticket was in it told you nothing you did not know from knowing it was a ticket. It had survived two previous reviews, because it is large and because filing a ticket there is never wrong.

Reasons with no category of their own

Address validation rejected a good address211held inside order sync
Return label request failed192held inside label printing
Carrier pickup not collected161held inside tracking

Five of seven passed every gate. Both failures are instructions rather than deletions, and they are different instructions: one category is retired, one is split.

Every guide on this recommends the same five categories. Every queue running them looks the same: one enormous bucket, one called Other, and the rest mixing what the customer said with what you would have to change. The help desk will not rescue that list. Zendesk's automatic tagging scans each ticket description for words longer than two characters and adds the top three matches against tags already in use, so it reproduces the taxonomy you already had. It cannot surface a category you have no tag for.

This pack derives the candidates from the text of the tickets you received, and lets the data choose how many there are rather than how many fit on a slide. Every candidate then goes through three gates. Enough volume to staff against. Two people independently filing the same ticket in it, against a floor of 0.60, below which methodological work on the agreement statistic treats agreement as inadequate. And a root cause mix that differs from the whole queue's, because a category whose mix is the queue's tells you nothing new.

Havenlea sells shipping and fulfilment software to online retailers. Its inherited five categories put 71.4% of six months of tickets into Technical Issue, which held nine distinct customer problems and five distinct fixes. Seven candidates came out of the derivation and five survived the gates. The third largest category in the queue, 358 tickets a month, failed two gates at once. Two coders agreed on it at 0.55, and its root cause mix sat 0.04 from the whole queue's. It was not a category. It was the absence of one, with a name.

The volume table that carries its own error bars

Volume is reported on both axes and with hours beside every count, because the two rankings disagree. Then the taxonomy is applied to a month it was never derived from, and the tickets it cannot place and the ones it places twice are published as rows.

Volume by Category, extract

Sorting by tickets puts the cheap problem first

Contact reasonWhat we would changeTicketsHoursBy ticketsBy hours
Setting up a shipping ruleDocumentation gap1,110254123
Tracking is not updatingCarrier API defect9921,41821
Label will not printCarrier API defect86092632
Rate at checkout is wrongCustomer configuration75242649
Setting up a shipping ruleCustomer configuration634282520
Label will not printHavenlea defect55384973

The largest cell by ticket count is 1,110 tickets and 254 hours. The largest by hours is 992 tickets and 1,418 hours, which is fewer tickets and more than five times the cost, and it ranks twenty-third by count. Crossing the two axes gives 0.108 normalised mutual information, so neither can be recovered from the other and a single column discards one of the two answers.

The held-out month

Both taxonomies applied to June, which neither had seen

TaxonomyPlaced in oneMatched twoPlaced in none
Ten categories, one per contact reason2,269  83.1%342  12.5%120  4.4%
Seven categories, derived and gated2,291  83.9%320  11.7%120  4.4%

The margin is small and it is not the argument

Both place about the same share of the month and both leave the same 120 tickets unplaceable. The argument is that all three numbers exist at all, for both lists, against tickets neither had seen. A double match is a ticket whose category depends on who picked it up. An unplaced ticket is one the taxonomy has no answer for. A published category list reports neither, so a reader cannot tell how much of its volume table is a coin flip.

The 320 double matches concentrate rather than scatter: 246 of them sit between the same two categories, so one tie-break rule clears three quarters of the problem.

What's in the pack

01

Category Register

The sheet the taxonomy is decided on, and every later sheet inherits its category list. One row per candidate: the cluster it came from, its highest weighted terms, tickets a month, and its largest root cause. Then the two measurements it had to earn, the agreement figure two independent coders produced and its distance from the queue's own root cause mix, and a verdict. Reasons that got no category of their own are on it too, marked absorbed with the host category and the share of it they represent. A reason the derivation could not separate is a finding rather than an omission.

02

Volume by Category

The reporting deliverable, on both axes rather than one, with handling hours beside every ticket count and both rankings shown. The two rankings disagree by twenty-two places at the top of the worked example, so a table sorted by count keeps the expensive problem below the fold. The owner sits on the root cause rather than the contact reason, since that is the thing being changed. The metric definitions pack does the same pinning job for every other number in the business. Once a category is ranked by hours, ticket deflection and content gap analysis checks which of them a help center article could actually shrink.

03

Untagged Queue

The tickets the taxonomy could not place and the ones it placed in two categories at once, with the fit score, the derivation floor it fell short of, and the margin between its top two. A permanent artefact rather than a clean-up list, because it is what the next revision is derived from. Read it by group and not by count: a concentration in one group means a category is missing, and a flat spread means the placement floor is too tight.

04

Trend

Month by month volume per category across the derivation window, with the held-out month beside the mean and the swing across months on the same row. It exists so a category that only exists because of one month's incident is visible as one, and so a category quietly disappearing from the queue gets retired instead of kept for continuity. The incident that created the spike is easier to trace against the account-level record the customer incident communication template keeps.

05

Taxonomy Definitions

What each category means, written so somebody holding a ticket can decide. Every definition carries the two figures it had to earn, and a category with no figures next to it is not in force. It also records what was retired and why, in enough detail that the argument for keeping a large useless category only has to be had once.

06

Tagging Guidance

When to set each field, why you categorise from the first inbound message rather than the thread, and the tie-break rules for the pairs that genuinely overlap, each written as a question about the ticket a person can answer. Also the four separate reasons the tag field is not the category field, which is worth reading before anybody reports off tags. Pairs with the naming convention pack when the fix is the namespace rather than the taxonomy.

07

Review Procedure

What happens monthly, quarterly and annually: read the unplaced by group, recompute the three gates, and once a year re-derive from scratch with no reference to the taxonomy in force. Every revision has to reconcile volume before and after, because a revision where the total does not match is a revision that lost tickets. Its sibling support ticket export template covers getting the data out completely in the first place.

How it works

  1. 1

    Send the tickets, not the summary

    Subjects and bodies, or the first inbound message on each conversation, six months if you have it. River holds the most recent month back before reading anything else, because the taxonomy has to be tested on tickets it was not derived from. Your current category list is useful as the comparison, not as the starting point.

  2. 2

    Derive the candidates and sweep the count

    Term weighting, then drop the terms on more than half the tickets, which is where the greetings and your own product name go. Then sweep the number of categories and read the separation. The worked example had ten distinct contact reasons in the queue and seven separated best, and that gap is the useful part rather than a rounding problem.

  3. 3

    Run the three gates per category

    Volume, agreement between two independent coders, and distance from the queue's own root cause mix. Per category and not just overall, because an overall figure of 0.77 routinely contains a category sitting at 0.60 by itself. Every failure comes back as an instruction with the tickets it held reconciled somewhere accountable.

  4. 4

    Test it on the month it never saw

    Both thresholds derived from the training distribution rather than chosen, then three counts: placed in exactly one category, matched two, placed in none. The double matches get grouped by pair, so the largest pair becomes one tie-break rule rather than a few hundred individually contested tickets.

Frequently asked questions

Why not just use the five standard support categories?

Because the biggest of them is not a category. On the worked example Technical Issue held 71.4% of the queue, containing nine distinct customer problems and five distinct fixes, with handling times spread almost exactly like the whole queue's. Grouping by those five categories accounted for 4.3% of the variation in handling time.

How many categories should a support taxonomy have?

It is a question for your tickets, not for a guide. Sweep the count and read the separation across the range. Expect fewer than the number of problems you can name: three of the worked example's ten shared their vocabulary too closely for anyone reading a ticket to separate, so splitting them would have failed on the queue.

Can't Zendesk's automatic tagging do this for me?

No, and its documentation says why. It matches ticket words against tags already in use, so it reproduces the taxonomy you have and cannot surface one you lack. It also skips tickets agents raise inside the help desk and may not work outside English, so it produces a systematically untagged population.

What makes a category worth keeping?

Three things, measured before it goes into the definitions. Enough volume to staff against. Two people independently filing the same ticket in it. And a root cause mix that differs from the whole queue's, because a category whose mix is the queue's mix tells you nothing new. Five of seven candidates passed all three.

What do I do about the Other bucket?

Run both gates on it and publish the figures rather than arguing. On the worked example it was the third largest category at 358 tickets a month, agreed at 0.55 against a 0.60 floor, and sat 0.04 from the queue's own root cause mix. It was retired, and the residue moved to a visible sheet.

Should the same field hold the customer's problem and the root cause?

No. They are close to independent, at 0.108 normalised mutual information on the worked example, so one column carries one answer and discards the other. Report volume on the pair. The sales and marketing SLA pack makes the same argument about a handoff whose two sides get measured as one.

How do I know the taxonomy will hold up next month?

Apply it to a month it was not derived from and publish three counts: placed in exactly one category, matched two, placed in none. The worked example placed 83.9%, double-matched 11.7% and could not place 4.4%. A category list quoting none of those has never been tested on tickets it had not already seen.

Derive the categories from your own queue, then test them

Edit with AI