River
Y CombinatorBacked by Y Combinator

Product & DesignFree

User Research Synthesis From Transcripts

Every coded quote carries the participant who said it, so a theme's count is people rather than mentions, and the dissent stays in.

Start here

River reads the transcripts rather than asking you to summarize them. Every transcript goes in, however your platform exported them, and each is read end to end rather than skimmed for the memorable parts. What comes back is a coded corpus: one row per extract, carrying the participant, the timestamp, the code, and the theme it rolls up to. Themes are then reported with three separate counts, so prevalence is a number you chose rather than a word like most or many standing in for one.

Everything that ranks for this query teaches a counting rule and then leaves you to do the counting. One template states it as law, that three participants independently saying the same thing is an insight, and attaches a blank grid to it. The best hands-on walkthrough on the subject demonstrates its own one-third rule while admitting the dataset is fake, because producing the real mapping by hand is too expensive to do inside an article.

Built for the researcher with nine transcripts and a readout on Thursday, and for the product manager who ran the interviews because the team has no researcher. Reach for it when the round is finished and the analysis has not started. Unsolicited feedback out of a support queue or a voting board is a different pile, and that is customer feedback triage. A finding that survives the room usually becomes a requirement, which is what the PRD writer is for. The change it specifies comes back later as an experiment readout with real numbers.

Five mentions can be two different findings

The word mention hides a choice. Braun and Clarke's 2006 paper, which most synthesis is loosely based on, names three levels at which prevalence can be counted. They are per transcript, per different speaker who articulated it, and per separate occurrence across the set. In one-to-one interviews the first two collapse together, so the live choice is people against occurrences. Those two rank themes differently, and a finding written as most participants tells you which was used only by accident.

Nine interviews, two themes. Approvals stalling onboarding appears twenty-six times, and one participant accounts for seventeen of those. Nobody knowing who owns the queue appears eleven times, spread across eight of the nine people. Ranked by mentions, approvals wins by better than two to one. Ranked by participants, it loses three to eight. Both rankings come from the same transcripts, and which one you see depends on a decision nobody wrote down.

Then there is the person who disagreed. Participant four said the approvals step had never blocked them, which is the row a summary drops first. Hennink and colleagues found the first interview alone produced 53 percent of new codes, so a codebook built early and never checked back against everyone is already leaning toward whoever you spoke to first. Reporting the count and the contradiction together is what stops synthesis becoming confirmation. It is also the cheapest defence when somebody in the room disagrees.

How it works

  1. Hand over the transcripts

    All of them, however your platform exported them, one file at a time if that is all it offers.

  2. Code at quote level

    Every extract gets a code, the participant who said it, and where in the session it appears.

  3. Count the themes

    By participant, by transcript and by occurrence, with the contradicting extracts kept alongside.

  4. Write the readout

    Findings with their supporting quotes, the dissent, and a deck built from the same rows.

What you get

  • One row per coded extract, carrying the participant, the timestamp and the code
  • Each theme counted three ways: by participant, by transcript and by occurrence
  • Every quote traceable to the point in the transcript it was taken from
  • The dissenting view recorded against the theme it contradicts, rather than dropped as an outlier
  • Themes named for the concept that unifies them, not the topic they sit in
  • The readout deck built from the same coded rows as the findings document

Common questions

How do I know whether five people said it or one person said it five times?

That is the count the sheet exists to produce. Every extract carries the participant who said it, so a theme reports the number of people, the number of transcripts, and the number of separate occurrences as three columns. When they disagree, the disagreement is the finding, and the row that caused it is one click away. A usability test report applies the same reversal to task failures.

How do I know the AI did not make a quote up?

Because every quote is a row pointing at a location in a transcript you supplied, and you can open it. Nothing in the findings document appears without an extract behind it. Where a passage is ambiguous or the speaker labels are unreliable, it is flagged rather than tidied into a clean attribution. With the audio in hand, an interview transcript workup flags what reads fine and scored badly.

How many interviews before the themes stop being new?

Fewer than you would expect to hear them, more than you would expect to understand them. Hennink and colleagues found code saturation at nine interviews but meaning saturation between sixteen and twenty-four. The coded corpus lets you check your own set instead of borrowing theirs, by showing which interview each code first appeared in.

What is the difference between a theme and a topic summary?

A theme has a central concept holding it together; a topic summary is everything people said about an area. Braun and Clarke call experiences of Y and benefits of X classic topic-summary theme names. Onboarding is not a theme. Approvals are treated as a favour rather than a step is one, and a team can act on it.

Do I need a second person to code the data?

Usually not, and the folklore here is stronger than the literature. Formal agreement scoring is uncommon in published qualitative work, and there are named situations where it does not apply, such as when developing the codes is part of the analysis. What helps more is cheap review: every code points at the extract behind it, so a colleague can disagree with a specific row.

Does this replace Dovetail or my Miro board?

No, and it starts where they leave you. A repository stores sessions and tags them; the board is where a team argues about clusters. What is missing between the two is the coded table, built by reading every transcript rather than the memorable ones. Export one session at a time if that is all yours offers. A single standout interview can also become a case study.

Can I present findings from a small sample?

Yes, if you report the sample honestly rather than dressing it up. Six participants stated as six participants survives a room; six written as most users does not. Each finding carries its participant count and the view that contradicted it, which is what a sceptical stakeholder is probing for. The problem statement slide and the counted job statements both come from these rows.

User Research Synthesis From Transcripts

Fill in the form and your workspace opens with the work already underway.