River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Qualitative Document Analysis Template

Three documents and three sheets that appraise every document on four criteria before a word of its content gets coded.

Free download  ·  No account needed

Document Register

Three documents excluded, one per criterion

Doc IDTypeDispositionReason
SR-03Situation reportExcludedAuthenticity — marked draft, not for release
PR-05Press releaseExcludedRepresentativeness — duplicate of PR-04
FAQ-02FAQ snapshotExcludedCredibility — capture failed to render
SR-04Situation reportIncludedPasses all three gate criteria

Meaning is assessed for every included document and never excludes one. The other three criteria can, and each excluded exactly one document here.

A document corpus is assembled after the fact rather than collected in the moment, and nothing vouches for a document's place in it the way an interviewer vouches for who said what in a transcript. John Scott's four-criterion appraisal grid for documentary sources is built for exactly that gap: authenticity, credibility, representativeness and meaning, asked of every document before its content counts as data. Three of the four can exclude a document outright. The fourth, meaning, never does; it is interpreted rather than adjudicated, a distinction most coding guides skip past.

The pack is three documents and three sheets. Sampling Rationale and Coding Frame fix the document universe and the five codes before a single document is appraised. Document Register holds the appraisal, two dates per document rather than one, and the disposition with its reason. Coded Extract Register holds the verbatim excerpts behind every coded passage, and Frequency by Source Type compares how often each code appears across genres instead of pooling them into one count that would hide the gap between them.

The worked example is one health department's communications during a foodborne-illness investigation: eighteen documents assembled, fifteen appraised into the coded corpus. A press release's stated detection date is the date the department went public, seven days after its own internal record shows the cluster was first suspected. A FAQ page reverses a factual claim between two of its own capture dates, with nothing on the page marking that anything changed. Open the pack in River and the agent appraises alongside you, or download the sheets and run your own corpus through the same register.

Every sheet in the pack

The Document Register (a subset of eighteen), the Coded Extract Register, and Frequency by Source Type.

Document Register

Eight of the pack's eighteen rows, including all three exclusions. Every included row also carries a meaning note, not shown here for space.

Doc IDTypeStated dateRetrievedDispositionReason
SR-01Situation reportMarch 6March 6IncludedPasses all three gate criteria
SR-03Situation reportMarch 15March 15ExcludedAuthenticity. Marked “DRAFT — NOT FOR RELEASE,” released in error alongside the final version
SR-04Situation reportMarch 16March 16IncludedPasses all three gate criteria
PR-01Press releaseMarch 13March 13IncludedPasses all three gate criteria
PR-03Press releaseMarch 25March 25IncludedPasses all three gate criteria
PR-05Press releaseApril 3April 3ExcludedRepresentativeness. Verbatim repost of PR-04 under a second URL; would double-count one message as two
FAQ-01FAQ snapshotNot dated on the page itselfMarch 17IncludedPasses all three gate criteria
FAQ-02FAQ snapshotNot dated on the page itselfMarch 22ExcludedCredibility. Accordion body failed to render before capture; retrieved text is the collapsed shell only

A live page's own field says “not dated on the page itself” rather than sitting blank. That is the finding, not an oversight.

Coded Extract Register

The verbatim excerpts behind the register's most consequential cells, six of the pack's thirteen.

IDDocCodeVerbatim excerpt
X-01SR-01C3 Detection timeline“Surveillance flagged three cases sharing a soft-cheese exposure over the four days ending March 6; the responsible processor is not yet identified.”
X-04PR-03C3 Detection timeline“The department first identified this cluster of illnesses on March 13.”
X-12SR-06C3 Detection timeline“This closure memo confirms the cluster's onset window as March 2 through March 6.”
X-07FAQ-01C2 Consumer guidance“No additional cases have been identified since March 17.” (captured March 17)
X-08FAQ-03C4 Uncertainty ack.“Additional cases identified after March 17 are under investigation.” (captured April 2)
X-09SR-05C4 Uncertainty ack.“Not yet confirmed whether the two most recently reported cases share the same exposure as the earlier twelve.”

X-01 and X-12 agree on March 6. X-04, the only press release stating a timeline at all, says March 13 instead, seven days later.

Frequency by Source Type

All five codes, compared by genre rather than pooled into one total.

CodePress releases (n=7)Situation reports (n=5)FAQ snapshots (n=3)Total (n=15)
C1 Processor or product named63312
C2 Consumer guidance given6039
C3 Detection timeline stated1506
C4 Uncertainty acknowledged1315
C5 Corrective action described43310

Situation reports state a detection timeline in five of five documents; press releases in one of seven, and that one names the wrong date.

What's in the pack

01

Sampling Rationale

Fixes the document universe before a single file is appraised: the date range, the issuing body, which genres are in scope and which are deliberately excluded. A genre like a live web page gets sampled differently from a dated archive.

02

Coding Frame

Five codes, each with an include rule, an exclude rule and an anchor example, applied identically across every genre in the corpus. No reflexivity note and no codebook change log, because a document does not have a rapport with whoever reads it.

03

Document Register sheet

Every document appraised on authenticity, credibility, representativeness and meaning before its content counts, with two dates recorded separately and a disposition reason on every row. Where cases get charted against fixed dimensions instead of documents, a framework matrix is the comparable structure.

04

Coded Extract Register sheet

The verbatim excerpt behind every coded passage, tied back to its own document's ID. For coding a study set rather than a mixed corpus, the extraction table preserves original wording the same way.

05

Frequency by Source Type sheet

Every code counted separately by genre rather than pooled into one figure, so a pattern that is common in one source type and rare in another does not average out to a number that describes neither. Coding a set of interviews instead runs through a codebook built for transcripts.

06

A document is not coded until its own record is filled in

The standing space rule every prompt reads first. Authenticity, credibility and representativeness can each exclude a document; meaning never does, and gets a note regardless of outcome.

How to use it

  1. 1

    Open in River, or download it

    Open the pack in River and the agent appraises and codes documents with you, or download the three Word documents and three CSV sheets instantly, filled in with the worked example.

  2. 2

    Fix the universe and the frame first

    Sampling Rationale defines what corpus answers the question and what is deliberately out of scope. Coding Frame turns the research question into five codes before any document is read for content.

  3. 3

    Appraise before you code

    Every document gets authenticity, credibility and representativeness checked, in that order, before a word of it is coded. A document that fails one stops there, with the reason on record.

  4. 4

    Compare across genres, then write up

    Frequency by Source Type compares how often each code shows up per genre rather than pooling them. The same read-the-record-before-the-claim discipline runs through building a position map from public statements.

Frequently asked questions

Is this template free?

Yes. Download the whole pack as Word documents and CSV sheets with no signup and no credit card. Edit with AI is a separate, optional path that has the agent appraise and code the documents with you. The template library holds the rest of the packs.

Does this replace NVivo, ATLAS.ti or MAXQDA for document analysis?

No. Those tools store the codes you apply to a document and are good at it. None of them ships a required appraisal step, so a draft mixed into a records-request batch or a duplicate press release gets coded exactly like a document that earned its place, unless you build that check yourself first.

My documents are all one genre, not three. Does this still help?

Yes, and Frequency by Source Type becomes a single-column count rather than a comparison. The appraisal step matters just as much with one genre, since a duplicate or a draft can turn up in a single-source corpus as easily as a mixed one.

What counts as a document's date when a web page has none printed on it?

The date you retrieved or captured it, recorded honestly as a retrieval date rather than guessed at as a publish date. The register holds both fields for every document, and leaving the stated-date field blank with a note is more accurate than inventing one.

What if I am not sure a document should be excluded?

Appraise it anyway and write the specific concern into its meaning note rather than guessing at a disposition. A borderline document included with a clear caveat is more useful later than one excluded on a hunch nobody can check.

What does 'Edit with AI' actually do?

It creates a free River account and installs this exact pack as a private workspace. Send the documents in whatever form they arrive, and the agent appraises each one, codes it against the frame, and builds the genre comparison and narrative with you.

How is this different from a systematic-review extraction table?

A systematic review's extraction table pulls comparable fields from published studies that already passed peer review. This pack assumes no such filter exists yet. It appraises the document itself, not just what it reports, before any content is trusted enough to extract.

Find out which of your documents earned their place in the corpus

Send the documents, however they arrived. The register comes back with every one appraised, not just filed.

Edit with AI