River
Y CombinatorBacked by Y Combinator

Product & DesignFree

Usability Test Report With Severity Ratings

Every finding starts from the task that failed and the participants it blocked, so the severity column is a count before it is a judgement.

Start here

River reads the sessions rather than asking you to grade them. Recordings, transcripts and the task log all go in, and every finding comes back attached to the task it broke, the participants who hit it, and the point in the session where it happened. Severity is computed before it is judged: how many people attempted the task, how many failed it, and how many recovered without help. The researcher's own rating sits beside that count as a separate column rather than replacing it.

The 0 to 4 severity scale every report template copies came from Jakob Nielsen. The same page that defines it warns that severity ratings from a single evaluator are too unreliable to be trusted, and recommends averaging three. Almost no team has three evaluators, so the number that decides which problems get scheduled is one person's estimate formatted as data. The templates ranking for this query hand you that column and leave you to fill it.

Built for the researcher writing up a moderated round on a deadline, and for the designer who ran the sessions because there is no researcher. Reach for it when the last session has finished and the findings document has not started. Interviews without tasks are a different analysis, and that is research synthesis. Unsolicited complaints from a support queue are customer feedback triage. Once a fix ships, whether it worked is an experiment readout.

A researcher reviewing session notes and task results on a laptop during usability analysis
Built for the write-up that happens after the last session, when the recordings are in and the findings document is not.

What a severity rating is actually measuring

Task completion is the one usability measure that does not require an opinion. A participant either finished the task or did not, so the result is a 1 or a 0 and the rate is arithmetic. Jeff Sauro's analysis of 1,189 tasks across 115 usability tests puts the median completion rate at 78 percent, which is the benchmark a finding gets measured against. A task your participants completed 40 percent of the time is not a matter of interpretation.

Six participants, one checkout flow, four findings. The label everyone squinted at got rated a major problem because it was the thing the researcher noticed most; it cost four people about ten seconds each and all six finished. The billing address field that silently rejected a valid postcode was rated minor, and it stopped five of six from completing at all. Ranked by rating, the label leads. Ranked by who failed the task, it is last.

The other half of getting a fix scheduled is making the finding watchable. A row saying users were confused by the billing step is an argument. A row carrying the participant, the timestamp and the forty seconds where they typed the postcode three times is not. Every finding here points at its moment in the session you supplied, so the person who has to fix it can watch it fail rather than take your word for it.

How it works

  1. Hand over the sessions

    Recordings, transcripts, notes and the task log, in whatever your research platform exported them as.

  2. Score the tasks

    Each participant marked complete or failed per task, with the ones who recovered without help noted.

  3. Derive the severity

    Findings ranked by who they actually blocked, with your own rating alongside and the gaps called out.

  4. Write the report

    Findings by severity with the clip reference on each, the issue register, and the readout deck.

What you get

  • One row per finding, carrying the task it broke and the participants who hit it
  • Severity derived from attempts, failures and unaided recoveries before any rating is applied
  • Your own severity rating kept as a separate column, with the disagreements flagged
  • Every finding pointing at the participant and timestamp where it happened in the session
  • Task completion rate per task, against the 78 percent median from published benchmarks
  • The readout deck built from the same rows as the findings document

Common questions

How is severity worked out if I only had one researcher?

From the sessions rather than from you. Each finding carries how many participants attempted the task, how many failed it, and how many got through only with help. That produces a rank before anyone rates anything. Your own rating still goes in its own column, and where the two disagree the report says so.

Is five participants enough to report on?

Five is enough to find problems and not enough to measure rates precisely, and the report treats those as different claims. A finding that blocked four of five participants is reported as four of five, never as 80 percent of users. Nielsen's own five-user recommendation is about discovery, not about producing a completion rate you can quote.

Can it work from recordings if I have no notes?

Yes. Recordings and transcripts are the primary input and the notes are a convenience. Where the task log is missing, completion is read from the session itself and each judgement call is flagged as inferred rather than recorded. Sessions where the participant went off-script are marked instead of being scored as a clean pass or fail.

Hotjar and FullStory do not export the video. Does that matter?

Not for a moderated study, where you have your own recordings. For unmoderated tools that only export metadata and heatmaps, the report works from the event and selector data and returns a shortlist of sessions worth watching yourself. The finding still names the screen element rather than the page.

What does the engineering team actually receive?

An issue register, one row per finding, with the task, the failure count, the derived severity, the screen element and the timestamp to watch. It drops into a backlog without rewriting. For the requirement wording that follows a finding into a spec, the PRD writer picks it up from there.

Does this replace my research repository?

No, it starts where the repository stops. Dovetail and Lookback store and tag sessions; neither produces the ranked findings document. Export one session at a time if that is all yours offers, and hand over the whole set. The coded output is a sheet you keep, not a view inside another tool. Filing the finished report by the questions it answered is what the research repository pack does.

Usability Test Report With Severity Ratings

Fill in the form and your workspace opens with the work already underway.