River
Y CombinatorBacked by Y Combinator

Research & PolicyFree

Interview Transcript Prep for Analysis

One speaker label per person across the whole corpus, every identifying mention replaced from a register, and the key in its own sheet.

Start here

River reads every transcript, inventories each distinct speaker string it finds, and maps them to actual people. Then it enumerates every identifying mention in the corpus with all the surface forms it takes, assigns a pseudonym to each real entity, and applies the replacements from that register rather than from a search box. You get the cleaned transcripts, one label per person, and the pseudonym key as a separate sheet you keep and never share. Two participants who consented to be named stay named.

Page one here sells transcription accuracy or lists accepted file formats. Neither addresses the two jobs that actually block coding, and both of those jobs turn on exact strings. A QDA importer builds one code per speaker, and MAXQDA states that the name of the participant must appear at the beginning of the paragraph, followed by a colon, with upper and lowercase taken into account. Anonymisation is a rewrite of those same strings. Doing one before the other breaks whichever comes second.

Built for the researcher whose transcription service changed format halfway through, the doctoral candidate whose ethics approval promised pseudonyms, and anyone about to import fourteen files and hope. Run it before the qualitative coding pass starts, because a label fixed after import does not retroactively split a merged code. The theme write-up needs one label per participant to count anything, the review synthesis writer handles the literature side, and more research space packs sit alongside this one.

Fifteen people, eleven speaker codes

An illustrative corpus of fourteen pharmacist interviews contained eleven distinct speaker strings. Four of them were the same interviewer, written as Interviewer, INTERVIEWER, I and Researcher across different files. One string, Participant, was five different pharmacists. Another, Speaker 2, was four more. So fifteen people arrive in the importer as eleven codes, and nine of the fourteen participants cannot be told apart at all. That is the count a thematic analysis needs, gone before coding starts.

The same corpus held two hundred and eighteen identifying mentions. A find-and-replace on full names catches one hundred and fifty-five of them and leaves sixty-three: thirty-one first names standing alone, fourteen bare surnames, nine sets of initials, nine possessives and transcriber misspellings. Nearly three in ten survive the search box. Worse, one participant and a patient's daughter were both called Sarah, so a global replace makes the participant describe herself in the third person.

So the register comes first and the replacement second, and the register lives in its own sheet. That is not a preference. Pseudonymisation is defined as processing where data cannot be attributed to a person without the use of additional information, provided that such additional information is kept separately. A key sitting inside the project file you email your supervisor is not kept separately. The UK Data Service adds that anonymisation should not be treated as a single technical step carried out at the end of a project.

How it works

  1. Send the transcripts

    Whatever state they are in, plus where they have to import and what must stay identifiable.

  2. Inventory the speakers

    Every distinct speaker string in the corpus, mapped to the person it actually refers to.

  3. Build the register

    Each real entity, its pseudonym, and every surface form it appears as, before anything is replaced.

  4. Apply and check

    Replacements run from the register, then a sweep for anything the register missed, reported to you.

What you get

  • Cleaned transcripts as Docs, with one speaker label per person across the whole corpus
  • The participant register as a Sheet, holding the pseudonym key and nothing else
  • Every identifying mention enumerated with each surface form it takes, before any replacing happens
  • Speaker labels in the exact form your QDA tool needs, colon and casing included
  • Name collisions flagged rather than merged, so two Sarahs stay two people
  • Timestamps normalised to one convention, or stripped, whichever your coding actually uses
  • A short note on what was removed and what deliberately stayed, for your ethics file

Common questions

Why not just anonymise inside my QDA tool after importing?

Because the mapping then lives inside the project file, which is the file you send your supervisor, your second coder and your co-authors. Pseudonymisation only holds when the key is kept separately, and a key stored in the same artifact as the data is not separate. Fix the strings before import, not after.

Can it work from recordings, or does it need transcripts?

Either. Send audio and it transcribes first, with the speaker labels already in the form your tool expects, so the preparation pass and the transcription are the same pass. Send finished transcripts and it works on those. Mixed sets are the normal case and it handles them together.

How does it know a name is a name rather than an ordinary word?

It reads for role and context, not for a name list. A pharmacy called Bridgeway, a colleague called Chris and a town called Maiden Vale all get proposed as entities with a reason attached, and you approve or reject each one in the register before a single replacement runs.

What about verbatim conventions, hesitations and overlapping speech?

Preserved by default, because whether you keep the ums is a methodological decision and not a cleanup one. Tell it you want intelligent verbatim and it removes fillers and repairs, logging what changed. It never silently smooths a quote you might later put in a results section.

Two of my participants consented to be named. Does that break it?

No, and it is worth stating plainly in the intake. Those two get their real names in the register with a note that consent covers attribution, and their transcripts stay identified. Everyone else in their transcripts still gets a pseudonym, including colleagues and patients who consented to nothing.

My timestamps are on every single line. Can they go?

Yes, or they can be reduced to one every few minutes, or kept as they are. Line-level timestamps from automatic transcription break paragraph-based speaker detection in most importers, so the default is to normalise them to the convention your chosen tool reads and say what was dropped.

Does the pseudonym key ever go into the transcripts?

Never. The register is a separate sheet, and the note that ships with the transcripts describes categories rather than individuals: what kinds of identifier were replaced, what deliberately stayed, and why. The only artifact that can re-identify anybody is the one you keep.

Interview Transcript Prep for Analysis

Fill in the form and your workspace opens with the work already underway.