River
Y CombinatorBacked by Y Combinator

Writing & MediaFree

Interview Transcript to Article Workup

Every quote comes back with its timecode, and the passages the model was least sure of become a relisten queue, not a clean sentence.

Start here

River works from the recording rather than from a transcript somebody already cleaned up. Every quote it keeps comes back with the timecode it was spoken at. Every passage the decoder itself would not vouch for comes back as a relisten queue with a duration attached. You get a quote bank you can check in a Sheet and the transcript organized by theme in a Doc. What you do not get is a draft, because the piece is the part that has to be yours.

The pages ranking for this term hand you a workflow that ends in writing the article, and so does the answer Google prints above them. None of them mentions that the transcript can be confidently wrong in exactly the places a reporter quotes from. That gap matters. A machine writing a smooth sentence over four seconds of unclear audio is how a misquote reaches print, and once it sits inside quotation marks the repair is a correction.

Built for reporters working from a recording, producers cutting an interview to tape, and anyone who must say where a quote came from. Reach for it once interview preparation has done the pre-interview homework, and before the first draft. If what you hold is text rather than audio, research synthesis codes transcripts across a whole study instead, and conflicting report reconciliation settles two numbers that will not agree. Where a story rests on documents rather than people, the FOIA request pack files the records request and counts its statutory deadline.

The accuracy number is not measured on the words you quote

Read the vendor's number and then read the footnote. On read audiobooks the same model scores 2.7 percent, and on the 231 recorded interviews in the CORAAL corpus it scores 16.2 percent. Six times worse on the same model, and your tape is the second one. Those are sociolinguistic interviews rather than press interviews, so treat 16.2 percent as the closest published stand-in and not as your own result. The direction is the point: fluent read speech is the easy case and a recorded conversation is not.

Averages also hide where the errors land. Across seven systems on real recorded earnings calls, person names came back wrong 42.1 to 51.7 percent of the time. The same systems reported overall rates of 11.3 to 17.8 percent. Names fail at roughly three times the headline rate, and every one of those systems is worse on names than on ordinary words. A reporter's exposure is concentrated in names, titles, places and numbers, which is the small subset the average is busy diluting.

Worse, one score cannot rank the damage. Drop a filler and an adverb from a 23-word sentence and you score 8.70 percent. Change the budget and the body that approved it and you score 8.70 percent again. The costs are not comparable, and AP's rule closes the escape route: quotations are not altered, even to fix word usage. A quote too murky to paraphrase should not run, and a confident transcription removes the murk from view, so that test never fires.

How it works

  1. Hand over the tape

    The recording as MP3, WAV, M4A or MP4, plus whatever rough transcript you already have.

  2. It reads and scores

    Every segment gets the decoder's own confidence reading alongside the words it produced for that stretch.

  3. Doubt becomes a queue

    Low-confidence stretches turn into timecodes with durations, rather than into smooth sentences you cannot audit.

  4. You check, you write

    Play the flagged seconds, correct what you actually hear, then write the piece yourself.

What you get

  • A quote bank as a Sheet, one row per quote, with speaker, topic and timecode
  • A relisten queue holding only the segments the decoder would not vouch for, with durations
  • Names, titles, places and figures pulled into their own list, because those break first
  • Every disagreement between your rough transcript and the recording, flagged as a check
  • The transcript organized by theme in a Doc, with speaker attribution on every passage
  • No draft, no suggested opening, no angle. The piece stays yours to write.

Common questions

Does it write the article for me?

No, and that is deliberate. The whole argument here is that a machine writing confident sentences over audio it could not resolve is how a misquote gets printed. Offering to guess at your opening would be the same failure one layer up. You get the quotes, the timecodes and the themes. The sentences are yours.

What does a low-confidence flag actually mean?

It means the decoder reported poor odds for that stretch of audio, not for one word. Readings arrive per segment, and a segment runs about twenty words, so a flag points at a passage and a timecode rather than at a syllable. That is still enough to send you to the right four seconds of tape.

What if my transcription tool reports no confidence at all?

Then flagging degrades to what the audio and the text jointly support, and the run says so rather than implying a signal it never received. Several newer models return a plain response with no per-segment scores at all. Where that happens you get flags from disagreements, crosstalk and the name and number checks instead.

Is there a published rule for spotting an unclear passage?

There is, and as written it cannot fire. The API reference says to treat a segment as silent once its no-speech probability rises above 1.0, which a probability never does. The same project's reference decoder ships 0.6 for that threshold. It is a documentation defect rather than a model one, but implement the docs literally and you get a detector that flags nothing, ever.

Can it tell me who said what?

It carries the speaker labels through and flags where they contradict themselves, which is the honest version. Who spoke is a separate guess from what was said, scored separately, and no word error rate contains any of it. Where two voices overlap you get the passage marked for a relisten instead of a confident attribution.

I already have a transcript. Is the audio really needed?

It works without it, but the best flags disappear. A transcript can only show you the problems it already admits: a garbled line, an existing inaudible marker, speaker labels that contradict each other. The passages worth checking are the ones that read perfectly. For text-only work, research synthesis is the better fit.

How big does the relisten queue get?

It scales with how much the run could not resolve, not with the length of the tape. Flag six percent of an hour's segments and you are relistening to about three and a half minutes rather than fifty-eight. Clean audio produces a short queue. A car park and a speakerphone produce a long one, which is itself worth knowing.

Interview Transcript to Article Workup

Fill in the form and your workspace opens with the work already underway.