River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Transaction Categorization Rules Template

Four documents and five sheets, where the rules that coded the month get measured against the policy instead of just applied.

Free download  ·  No account needed

Rule Performance

[Entity] — the month ending [__ ___ 20__]

Coverage is the share the rules coded with nobody in the loop. The error rate is the share of those codings a re-derivation from the policy disagreed with. Neither number is ever reported without the other.

MonthTxnsAutoReviewQueueCoverageWrongError rateError $

When one merchant carries two categories, it is one of exactly three things

It sells on both sides of a stated dollar thresholdAmount band on the rule
The policy changed on a day you can nameRetarget the rule, restate the earlier months
It genuinely sells across categories nothing separatesNever automate it, and say so

Coverage rising while the error rate falls is a rule set learning. Coverage rising on its own is a rule set getting more confident, which is a different thing and usually worse.

Every categorization template you can download is a list of category names, and every categorization tool reports one number about how it used them: coverage, the share of transactions it coded without asking. None of them report the error rate. The two can move in opposite directions, and the gap between them is where a P&L quietly stops meaning the same thing twice, because the coding nobody reviewed is the coding that was automatic.

This pack keeps the rules and measures them. Yarnell Mechanical, an illustrative contractor, closed February at 88.9 percent coverage with five transactions queued, which reads as finished. Re-deriving the month from the written policy found 38 codings that disagreed, worth 9,742, and 32 of those had fired automatically so nobody had read them. The largest, 5,549, came from a merchant the rules were right about twenty-seven times out of thirty. Feed it from converted statement PDFs or a reconciled month.

Two of the categories are decided by amount rather than by vendor, and both come from a published rule. A taxpayer with no applicable financial statement may expense tangible property up to 2,500 per invoice or item, so a tool distributor selling on both sides of that line needs an amount band rather than a category. And only half of a food or beverage expense is allowable, which is why meals buried in job materials claim twice what they should. Run it beside the monthly close pack.

Four of the sheets, filled in for one quarter

Rule Performance, where February went wrong, the three causes of a split merchant, and the P&L with recoding split out.

Rule Performance

Yarnell Mechanical, Inc., January to March 2026. An illustrative plumbing and HVAC contractor. 914 transactions, 425,168.

MonthTxnsTotalAutoReviewQueueCoverageBy $WrongError rateError $Rules
January311150,823164876052.7%38.7%239.2%5,20260
February305139,19027129588.9%82.1%3812.7%9,74265
March298135,15526924590.3%88.6%00.0%070

Between February and March coverage rose 1.4 points and the error rate fell 12.7. A report publishing only the first number would have called those two months indistinguishable. January's 52.7% is not a bad month, it is what a rule set with no history produces: the only way to raise it on day one is to promote rules on a single observation, which is how a rule set learns a coincidence.

Where February went wrong

38 codings, 9,742, across twelve merchants. Sorted by dollars, because sorting by count buries the answer.

MerchantWrongDollarsTierRule saidPolicy says
Ridgid Direct25,549AutomaticSmall tools and equipmentEquipment
ServiceTitan11,180ReviewOffice suppliesSoftware and subscriptions
Amazon Marketplace7661AutomaticOffice suppliesJob materials, small tools
Panera Bread7508AutomaticJob materialsMeals
Casa Moreno Grill5482AutomaticJob materialsMeals
Subway8412AutomaticJob materialsMeals
Buckeye Diner3351AutomaticJob materialsMeals
QuickBooks Online1189ReviewOffice suppliesSoftware and subscriptions
Fleetio1168ReviewOffice suppliesSoftware and subscriptions
Google Workspace1144ReviewOffice suppliesSoftware and subscriptions
Dropbox172ReviewOffice suppliesSoftware and subscriptions
Adobe Acrobat124ReviewOffice suppliesSoftware and subscriptions
TOTAL389,74232 of the 38, worth 7,965, were automatic and read by nobody

Ridgid Direct is the row that pays for the sheet. Sorted by count it is eighth and easy to skip; sorted by dollars it is first, and it is the only row that changes the balance sheet rather than the P&L. The six subscriptions underneath it were reviewed by a person every single time and confirmed wrong twelve times running, because review catches a rule that is wrong in an unfamiliar way and never one that is wrong in a familiar way.

Judgement Notes

Twelve merchants carried more than one category. Three causes, three different fixes, and picking the wrong one is what makes a category mean two things.

MerchantCategories carriedCauseFixEffective
Ridgid DirectSmall tools and equipment / EquipmentSells on both sides of a stated thresholdAmount band at 2,5002026-02-01
Casa Moreno GrillJob materials / MealsPolicy changed on a dated dayRetarget and restate2026-02-01
Panera BreadJob materials / MealsPolicy changed on a dated dayRetarget and restate2026-02-01
SubwayJob materials / MealsPolicy changed on a dated dayRetarget and restate2026-02-01
Buckeye DinerJob materials / MealsPolicy changed on a dated dayRetarget and restate2026-02-01
ServiceTitanOffice supplies / SoftwarePolicy changed on a dated dayRetarget and restate2026-02-01
QuickBooks OnlineOffice supplies / SoftwarePolicy changed on a dated dayRetarget and restate2026-02-01
Google WorkspaceOffice supplies / SoftwarePolicy changed on a dated dayRetarget and restate2026-02-01
FleetioOffice supplies / SoftwarePolicy changed on a dated dayRetarget and restate2026-02-01
DropboxOffice supplies / SoftwarePolicy changed on a dated dayRetarget and restate2026-02-01
Adobe AcrobatOffice supplies / SoftwarePolicy changed on a dated dayRetarget and restate2026-02-01
Amazon MarketplaceJob materials / Office supplies / Small tools / UniformsGenuinely sells across four categoriesNever automate2026-02-01

Ridgid Direct, all thirty transactions

PopulationCountAmount rangeCategoryWhat a flat rule did
Consumable blades and fittings24113 to 332Small tools and equipmentRight 24 times
Press tools62,720 to 3,457EquipmentWrong 3 times, 9,006 expensed

The rule promoted on the frequent answer, because eight of Ridgid's first nine transactions were consumables. The threshold is 2,500 per invoice or item and the merchant sits astride it, so the merchant name was never the deciding fact. Amazon is the opposite case: 48 transactions, 4,306, in four categories at overlapping amounts, so it is marked never to automate at a disclosed cost of about fourteen transactions and 1,217 a month reaching a person.

Monthly P&L

January restated onto the corrected rules. Recoding plus activity equals the reported change on every row, and the recoding column sums to zero or a transaction is missing.

CategoryJan as reportedRecodingJan restatedActivityMarchReported changeRead the change as
Job materials52,536(2,106)50,429(3,140)47,290(5,246)Activity
Meals02,1062,106(127)1,9791,979Recoding is the larger half
Office supplies3,635(1,777)1,8582042,063(1,573)Recoding is the larger half
Software and subscriptions01,7771,77701,7771,777Recoding is all of it
Small tools and equipment13,031(3,457)9,574(1,674)7,901(5,130)Recoding is the larger half
Equipment3,1503,4576,607(840)5,7672,617Recoding is the larger half
Thirteen other categories78,471078,471(10,093)68,378(10,093)Activity
TOTAL150,8230150,823(15,668)135,155(15,668)Recoding must be zero

Gross recoded for the quarter is 7,340, which is the sum of the absolute values halved. The net is always zero and says nothing; the gross says how much of the P&L changed categories without anything being bought. Software did not appear this quarter: Yarnell spent the same 1,777 in January, filed under office supplies, and activity on that row is exactly zero. Job materials looks 10.0 percent down and is 6.2 percent down once the meals leave it. And the 3,457 moving into Equipment is a press tool that was expensed above the capitalization threshold, which changes the balance sheet rather than the P&L.

What's in the pack

01

Rule Performance sheet

One row per month carrying coverage, the queue rate and the error rate, each by count and by dollars, plus how many of the wrong codings were automatic. Coverage never appears in this pack without the error rate beside it.

02

Merchant Rules sheet

The rule set as it stands, with the category, the tier, how many times the merchant has been seen, whether the rule carries an amount band, and whether it is marked never to automate.

03

Category Map sheet

Per category: what belongs, what does not, the test that decides a hard case, and the tax treatment that makes the line worth drawing. Two rows carry an amount threshold rather than a judgement.

04

Uncategorized Queue sheet

Every transaction that reached a person because no rule covered it, with the raw bank descriptor, the merchant key, why it queued and what was decided. The only place the system tells you what it does not know.

05

Monthly P&L sheet

Each category's change split into recoding and activity, with the earlier month shown as reported and restated. The recoding column sums to zero across all categories or a transaction is missing.

06

Categorization Policy

The dated policy, the date it came into force, what was restated onto it, and the two amount tests that are not house style. Written so the answer to a hard case stops depending on who is looking.

07

Judgement Notes

A row per merchant that carried two categories, with the cause named as one of exactly three and the fix that follows. The three look identical in the data and have nothing in common.

08

Rule Set Report

The three months side by side, the wrong codings grouped by merchant and sorted by dollars, what the review changed, and the three signals worth watching next month.

09

Reading the Bridge

Five conclusions the recoding split stops you from reaching, each worked through in full, including the category that appears to have grown from nothing and did not.

10

Space rule

Coverage is not accuracy, and a merchant is not a category. It governs every prompt here, and it is why nothing in this space will raise a coverage number by promoting a rule faster.

How to use it

  1. 1

    Open in River, or download it

    Open the pack in River and let the agent build it from your own transactions, or download the blank Word and CSV files instantly with no account.

  2. 2

    Send two or three months of transactions

    Bank and card, plus whatever category list and coding rules you already run. Two months is the minimum worth analyzing, because a rule that has never been tested against a second month cannot be measured at all.

  3. 3

    Run the consistency test before writing any rules

    Collapse the raw descriptors down to merchant keys, then ask which merchants were already coded two ways in the history you have. That list needs no policy and no rules, and it is the most useful thing on the table.

  4. 4

    Code the month, then measure it

    Re-derive every category from the policy independently of what the rules said, and report the error rate beside the coverage figure with the wrong codings sorted by dollars.

Frequently asked questions

Is this template free?

Completely. Take the Word documents and CSV sheets with no signup, no card and no trial. Edit with AI is an optional second path for anyone who would rather the agent code and measure their own transactions than fill the grid by hand. More packs sit in the template library.

Why measure the rules instead of just improving them?

Because you cannot tell the difference without measuring. A rule set that codes nine transactions in ten unattended looks identical, from the coverage number alone, whether it is right on all of them or wrong on all of them. Promotion is the moment a rule stops being read, which is exactly when a wrong one becomes expensive.

What makes one merchant carry two categories?

Exactly three things, and the fixes are unrelated. It sells across a stated dollar threshold, so the rule needs an amount band. The policy changed on a nameable day, so retarget it and restate the earlier months. Or it genuinely sells across categories nothing separates, so refuse to automate it and say so.

How is this different from the bank reconciliation pack?

Different question about the same transactions. The bank reconciliation pack asks whether the cash balance is real and dates what will not reconcile. This asks whether a category means the same thing twice. Card months with missing receipts route through the receipt matcher.

Does it separate personal spending from business spending?

Not directly, and deliberately. This space codes against a policy and measures whether it did so consistently. Deciding whether a charge belongs to the business at all is a different judgement with a tax consequence per category, handled in the owner draw review.

What format are the downloaded files?

Four .docx documents covering the policy, the judgement notes, the rule set report and the bridge, plus five .csv sheets, all in a single zip. Everything opens natively in Word, Pages, Google Docs, Excel, Numbers and Sheets with no conversion step.

What does Edit with AI actually do?

It spins up a free River account with this pack installed as a private workspace, already primed to collapse your descriptors into merchant keys, code each month against a written policy, and report coverage beside the error rate. Send nothing and it writes nothing.

Measure the rules, do not just run them

Download the blank pack as Word and CSV files, or open this exact pack in River and let the agent code and measure your own transactions.

Edit with AI