Lecture Caption and Transcript Template
Four documents and four sheets, built around a correction list scoped to the course rather than to each separate recording.
Free download · No account needed
Terminology Correction List
[Course code] — the terms this course keeps losing
One row per term, built for the course rather than for a recording, applied to every recording in it.
| Term | Machine output | Resolution | Pair partner | Context test | First seen |
|---|---|---|---|---|---|
| — | — | required | — | — | — |
| — | — | required | — | — | — |
| — | — | required | — | — | — |
| — | — | required | — | — | — |
| — | — | required | — | — | — |
| — | — | required | — | — | — |
Resolution is the one cell here that is never left blank. Substitution means the misheard form is not a word this course uses, so the row applies without reading around it. Context required means it is, the partner column names the other word, and the test replaces the substitution.
Every guide to captioning a lecture recording closes on the same line. Generate the automatic track, then review it carefully, paying particular attention to technical terms and proper nouns. The line is correct. It is also the entire job, restated as a reminder, and it carries no mechanism, so it gets done from memory once per recording. The vocabulary a course keeps losing is rediscovered every week by whoever happens to be doing that week's pass.
This pack scopes the correction list to the course rather than to the file. The worked audit in it covers BIOL 2140 Human Physiology at Brackenridge College: 24 recordings, 1,412 minutes, 1,043 logged corrections. 604 of those were course terminology, and they came from 37 distinct terms that were present in a recording 312 times. So 275 of the 312 encounters were a term somebody teaching that course had already decided how to spell.
Fourteen of the 37 terms are half of a bidirectional pair, where the form the recogniser produced is itself course vocabulary: homeostasis and hemostasis, afferent and efferent, neuron and nephron. Those 14 carry 256 corrections, and 17 recordings hold both members of a pair, so every row states a resolution class and pair rows carry a context test instead of a substitution. Terms come first in every pass, which is also why the rebuild of an inherited course starts from the same list.
What you get
Terminology Correction List
One row per term the recogniser gets wrong, with the form it produced, a resolution class, and the recordings it has appeared in. This is the artifact that outlives the course, and the reason the next term costs less than this one.
Caption Review Log
One row per correction pass, splitting what it found into course terminology, cue timing, numbers and units, and speaker labels. The split is what shows which part of the work repeats and which part is genuinely new each week.
Media Register
One row per recording: runtime, capture source, transcript word count, measured speech rate, the format published to each destination, and whether the slides carry content the audio never says.
Accessibility Compliance Check
One row per criterion, each with its own stated scope, because 24 recordings and 30 caption file deliveries are different denominators. Rows for SC 1.2.2, 1.2.4 and 1.2.5 sit next to encoding and destination-format checks, in a form an accreditation self-study can cite directly.
Caption Style Guide
Line length, cue duration, how student questions are attributed, how numbers and units are written, and the reading-rate ceiling. Settled once for the course so two reviewers produce the same file.
Caption Format Reference
What WebVTT and SubRip each actually require, written from the specifications rather than from a converter's help page, plus what survives a conversion between them and what does not.
Correction Procedure
The five passes in fixed order, terms first and timing last, because correcting terms moves where cues want to break and doing timing first means doing it twice.
Correction Findings
The full worked audit the sheets arrive shaped by, including the quarter-by-quarter split where terminology corrections more than double while new terms fall from 14 to 5.
How it works
- 1
Send the course, not just a file
The reading list, the slide decks, the glossary, a past exam paper, the syllabus. A course publishes its own vocabulary weeks before it says any of it out loud, so half the correction list is writable before anybody opens an audio file.
- 2
Send a machine transcript
The raw automatic output from one recording or several, uncleaned. River reads it against the vocabulary, proposes correction-list rows, and marks every row where the misheard form is itself a term the course uses.
- 3
Run a pass per recording
Terms, numbers, speaker labels, sense, then timing. Each pass fills a Caption Review Log row by category, so by the fourth recording the log is already saying which category is repeating.
- 4
Publish and check the course
One corrected cue list generates a file per destination, so a later move to another platform carries the corrections with it, and the compliance check states per criterion what is in scope and what is passing.
Frequently asked questions
What format are the downloaded files?
The four documents come as .docx and the four sheets as .csv, so they open in Word, Pages, Google Docs, Excel, Numbers and Sheets with no conversion step. The download needs no account. Caption files themselves are a separate output, covered by the Caption Format Reference.
What does Edit with AI actually do?
It installs this pack as a private space and primes River to fill it from what you send. In practice the first useful output is a draft correction list built from your reading list and slide decks, before a single recording has been reviewed.
Are automatic captions enough for accessibility?
Word accuracy is a different measure from conformance. WCAG 2.1 defines captions as a synchronized alternative for both speech and non-speech audio information, including speaker identification, and a speech recogniser produces neither, so those two start at zero in every recording however clean the transcript reads.
Should I publish WebVTT or SubRip?
Whatever the destination parses, and keep one corrected cue list that generates both. A WebVTT file begins with the literal string WEBVTT, is encoded UTF-8, and takes a full stop before the milliseconds; a SubRip entry needs a sequence number, a comma, and a blank line closing it.
Does this work for a course that is not physiology?
The worked example is physiology because the confusable pairs are vivid there, but the mechanism is subject-neutral. Any field with dense vocabulary produces the same shape: a small set of terms carrying most of the corrections, and a subset whose misheard form is also a real term the course design already names.
How fast is too fast to caption verbatim?
Section 508 guidance puts the practical ceiling at 180 words per minute, about three words a second, past which a captioner has to choose between cues that vanish before they can be read and cues that drift out of sync. The Media Register measures the rate per recording.
Does this cover live sessions?
Not directly, and the compliance check says so rather than hiding it. Live captioning is a service somebody books before the session starts, so it appears as its own row with an owner outside this space, next to the recorded rows it keeps getting confused with.
Find the terms your course keeps losing
Send a machine transcript, or just the reading list and the slide decks. River drafts the correction list first, then runs the passes against it.
Edit with AI