River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Internal Tooling Requirements Template

Internal tooling is discovered in the first week of incidents, because planning runs forwards from the feature. This template runs backwards, from the week after launch.

Free download  ·  No account needed

Every launch creates work that somebody has to do by hand afterwards, and almost none of it appears in the specification. Not because anyone is careless, but because specification runs forwards from the feature and the operational work sits downstream of it, past the point where requirements are still being written. So this pack starts the week after launch and walks backwards. Six workflows in the worked example produced 41 distinct actions a human would have to perform. The launch document named four of them.

The mechanism is one rule on every row: no fallback may end at asking an engineer without a volume and a duration. Those two multiply into minutes a month, the only form in which this argument has ever beaten launch scope. Google's SRE book named the category and gave it a ceiling, defining toil as the manual, repetitive, automatable work that scales linearly as a service grows and capping it at half of an engineer's time. A launch that creates toil without counting it has moved the cost, not avoided it.

Coverage and readiness are also asked separately, because they return different numbers. AWS built the same idea into its operational readiness reviews, which distil what actually went wrong into a curated question set rather than a checklist written from the design. Sits next to the non-functional requirements the same release needs and the constraint brief behind them. It is not a support runbook, which documents a process that already works.

Forty-one actions, eighteen gaps, one full-time engineer

Six workflows walked forward from go-live, every fallback costed, and the three builds that pay for themselves inside a quarter.

Workflow Impact Register

Illustrative, for a fictional B2B product called Padgett launching usage-based pricing. Six support and operations workflows walked forward from the day after go-live. A representative sample of the 41 rows.

WorkflowAction after launchWhoA monthTool todayVerdict
Billing disputeRead the invoice charges back to the customerSupport agent240Billing consoleReady
Billing disputeView the metering trace behind one usage lineSupport agent88Metering consoleTool exists, rehearsal failed
Billing disputeCorrect a mis-metered usage recordPlatform engineer96NoneAsk an engineer
Billing disputeShow which of the customer’s own users generated the usageNobody46NoneNo path at all
Plan changeApply a proration overrideBilling operations64Billing consoleTool exists, rehearsal failed
OverageConfirm current usage against the plan capSupport agent140NoneTwo screens and a subtraction
OverageRe-run a failed nightly usage rollup for one accountPlatform engineer28NoneAsk an engineer
RefundReverse a metered event after invoicingNobody31NoneNo path at all
6 workflows41 actions, 4 of them named in the launch document3 roles3,76423 covered, 18 gaps17 genuinely ready

Coverage and readiness are two questions and they returned different answers. 23 actions had a tool against their name. When the person who would actually perform each one was asked to do it on a real account, 6 could not: four were permissions nobody had thought to request, two followed runbooks describing a screen that had moved two releases earlier. Genuine readiness was 17 of 41 rather than 23 of 41, and every one of those six was a day of work rather than a build. They would otherwise have been found by a customer.

Tooling Gap List

The eight gaps whose fallback is an engineer with production access, sorted by cost. Volumes are steady state at month three. Durations were timed on a real case, including the audit note and the customer reply.

GapA monthMinutes eachHours a monthRank by volumeRank by cost
T-07 Usage breakdown behind a disputed invoice743543.221
T-03 Correct a mis-metered usage record962540.012
T-11 Re-run a failed usage rollup for one account284018.743
T-14 Move an account between pricing models175515.664
T-27 Recover a metering event dropped at ingest99013.585
T-22 Reconcile a partial credit against a usage line233011.556
T-16 Back-date a plan change12459.077
T-19 Clear a stuck overage notification48108.038
8 rows that end at ask an engineer30710 to 90159.4one full-time engineer, in nobody’s estimate

Two rankings, and where they disagree is the finding. T-19 is the third most frequent gap and the cheapest of the eight, because it takes ten minutes. T-11 is fourth most frequent and third most expensive, because it takes forty. Support intuition tracks frequency, since frequency is what a day feels like, and engineering capacity tracks duration. The row raised in every meeting is the one to build last. Seven further gaps land on support rather than engineering, at 115.1 support-hours a month, and the two totals are kept apart because they compete with different things.

What the 18 gaps actually cost

The decision rule was written before the list was sorted: build anything that pays back inside one quarter. On this list it selects exactly three rows.

DecisionRowsCost to buildRemoves each monthPayback
Build, into launch scope35.5 engineer-weeks101.8 engineer-hours2.16 months
Live with, engineer time58.5 if built57.6 engineer-hours5.90 months
Fix with copy, not a tool30.6 engineer-weeks67.2 support-hoursunder a month
One small support screen10.5 engineer-weeks18.7 support-hoursunder a month
Live with, support time30.029.2 support-hours absorbedn/a
No path, answer scripted3n/a245 requests a month refusedn/a
18 gaps186.6 recommended274.5 hours a month costed3 days to produce

The timing argument is the one that decides whether this is worth doing before launch. Found in the specification, the three builds ship with the launch and the escalation never happens. Found in week one of incidents, nobody is free for three weeks, and 5.5 engineer-weeks at realistic allocation lands the fix eleven weeks after go-live. The escalation paid while waiting is 258 engineer-hours, against a build worth 220. Waiting costs more than building did, and the register that finds it takes three days across a product manager, a support lead and an operations lead.

What is in the pack

01

A walk forward from the day after launch

Each workflow narrated from go-live rather than from the specification, stopping at every point where a human has to act. The unhappy paths produce most of the rows.

02

A fallback column that has to carry a number

Every answer that resolves to a person with production access gets a volume and a duration. A gap without those two is an opinion, and opinions lose to launch scope.

03

Coverage and readiness asked separately

Every covered action rehearsed once, on a real case, by the role who will perform it. Six of 23 failed in the worked example, and none of the fixes was a build.

04

Support time and engineer time kept apart

Both are real and they compete with different things. Support time competes with the queue. Engineer time competes with the roadmap, which is why it moves decisions.

05

A ranking by cost that disagrees with the one by volume

Frequency is what a day feels like, cost is frequency times duration. The gap raised in every meeting is usually the one to build last, and only the sort shows it.

06

A build list, a live-with list and a written no

Each rejection carries its cost and its reason, so it is not re-argued next quarter, and each refusal names the alternative that does exist.

How it works

  1. 1

    Send the launch and the workflows

    Ticket categories, escalation threads and recent incident reviews are worth more here than the process documentation.

  2. 2

    Walk each one forward

    Every point where a human acts becomes a row with a role, a frequency, a tool and a fallback.

  3. 3

    Rehearse what looks covered

    The person who will do it does it, on a real account, with nobody helping and a stopwatch running.

  4. 4

    Cost, rank, decide

    Volume times duration, sorted by cost, against a payback threshold written down before the list was sorted.

Frequently asked questions

What do I need before this is useful?

The launch itself, the support and operations workflows it touches, and whatever admin tooling exists today. Ticket categories and recent escalation threads are the most valuable input, because they are the only record of what people actually end up doing rather than what the process says they do.

Why walk backwards from launch instead of forwards from the feature?

Because forwards is how the work gets missed. Everything support and operations will have to do sits downstream of the feature, past the point where anybody is still writing requirements. In the worked example the walk found 41 actions against the four the launch document named.

Is a tool existing not the same as being ready?

No, and the gap between them is the most consistently surprising output here. Six of 23 covered actions could not be completed by the person who would perform them: four permissions nobody had requested, two runbooks describing a screen that had moved. Readiness was 17 of 41, not 23 of 41.

How do I estimate a volume I have never measured?

From the closest comparable launch where one exists, from the pilot where one does not, and from the support lead and operations lead agreeing a number out loud as a last resort. Unhappy paths run low in every organisation, so mark guesses as guesses and re-estimate after a month.

Does this recommend building everything?

The opposite, and a register that did would not be useful. In the worked example three of eight engineer-facing gaps are built, five are lived with for a stated reason, three support rows are answered with copy rather than a screen, and one request is refused in writing.

What about the things nobody can do at all?

They get recorded with the number of times a month they are going to be asked for, which is often the strongest roadmap input in the document. One of them here is asked 168 times a month. Another is a correct design decision rather than a gap, and it gets a refusal with a scripted answer.

How late is too late to run this?

After launch it still pays, and it costs more. The three builds in the worked example are worth 5.5 engineer-weeks. Started in week one of incidents they land eleven weeks after go-live, and the escalation paid while waiting comes to 258 engineer-hours against a 220-hour build. Retiring something later has the same shape.

Find out what your launch is about to hand to support

Send the launch and the workflows it touches. What comes back first is a count: how many actions the launch creates, and how many of them anyone can actually perform.

Find the tooling my launch needs