River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Intercoder Reliability Template for Research

Two documents and three sheets that score every code separately, flag whichever fall short, and hold each flagged code's disputed extracts in a register.

Free download  ·  No account needed

Agreement by Code

Five codes, scored separately against the same 120 excerpts

CodeRaw agreementKappaBandFlag
Described mentor check-ins as helpful90.8%0.815Almost perfectNo
Requested more classroom-management support90.0%0.100SlightYES
Described the program as effective overall70.8%0.391FairYES
Named a specific curriculum resource received97.5%0.908Almost perfectNo
Named a specific administrative barrier to implementation95.0%0.640SubstantialNo

Two codes sit within a point of each other on raw agreement, 90.8% and 90.0%. Their kappas are 0.815 and 0.100, eighteen bands apart.

A dual-coded qualitative study usually reports one number: a kappa next to the codebook, footnoted that disagreements were discussed and resolved. Pooling every code into that figure hides which code it punishes. Cohen's kappa corrects for the agreement two coders would reach by guessing, and that correction has a documented failure: a rare or skewed code can return a kappa near zero from a handful of ordinary calls. A published appraisal of 57 clinical trials found exactly this: 84.2 percent raw agreement produced a kappa of only 0.042, because one category covered nearly every trial.

The pack is two documents and three sheets. Reliability Procedure fixes the coders, the sample size, the statistic and the flag threshold before anyone codes anything, so a threshold never gets picked after seeing which codes look bad. Agreement by Code scores every code separately against the excerpts and flags whichever fall under that line, read against the standard bands rather than in isolation. Disagreement Register holds the disputed extracts behind every flagged code, quoted verbatim with why the coders split. Codebook Refinement Log records what changed, or that nothing needed to.

Reading the register is what tells a paradox apart from a real gap. In the worked example, one flagged code's twelve disagreements share no pattern, so two anchor examples get added and the definition stays put. A second flagged code's disagreements repeat one boundary problem six times in an eight-row sample, so an exclusion rule gets written instead, and Disagreement Resolution Note records which happened and why. Open the pack in River and the agent scores the codes with you, or download the sheets and read your own codebook's disagreements the same way.

Every sheet in the pack

Agreement by Code, the Disagreement Register, and the Codebook Refinement Log.

Agreement by Code

120 excerpts, two coders, five codes, scored separately rather than pooled into one figure.

CodeScoredBoth yesBoth noDisagreeRaw agr.KappaBandFlag
Described mentor check-ins as helpful12048611190.8%0.815Almost perfectNo
Requested more classroom-management support12011071290.0%0.100SlightYES
Described the program as effective overall12030553570.8%0.391FairYES
Named a specific curriculum resource received1201899397.5%0.908Almost perfectNo
Named a specific administrative barrier to implementation1206108695.0%0.640SubstantialNo

The flagged code with the lowest kappa, 0.100, is not the flagged code with the lowest raw agreement, 70.8%. Reading the register is what tells the two problems apart.

Disagreement Register

A sample of the excerpts behind the two flagged codes above, quoted verbatim with why the two coders' calls split.

IDCodeExtractJRTMPattern
S-006Classroom-management support“If there’d been a session on de-escalating things before a blow-up happens, not just after, I’d have gone.”YesNoNo shared pattern
S-077Classroom-management support“Where was the person who could’ve told me what to do with a kid who just won’t sit down? Anybody?”YesNoNo shared pattern
S-112Classroom-management support“If my mentor had come from a middle school like mine, half of this would have sorted itself out.”NoYesNo shared pattern
S-004Program effective overall“That one meeting where my mentor walked me through the whole grading system saved my first quarter.”YesNoSingle interaction generalized
S-052Program effective overall“The feedback on my October observation was the most useful thing anyone told me all year.”YesNoSingle interaction generalized
S-098Program effective overall“It’s fine. I’d do it again, I guess.”NoYesUnrelated: hedge weighted differently

Six of eight sampled disagreements on the second code shared this one shape. None of the twelve on the first code shared any.

Codebook Refinement Log

What changed after each flagged code's register got read, or the decision that nothing needed to.

CodeKappaRows readChangeRationale
Classroom-management support0.10012 of 12Anchors added, definition unchangedAll 12 disagreements were independent, defensible calls; no two shared a reason
Program effective overall0.3918 of 35 (sampled)Exclusion rule added6 of 8 sampled disagreements shared one boundary problem: praise for one interaction generalized to the whole program
Program effective overallPending rescore0 of remaining 27Rescore scheduledThe remaining 27 disagreements have not yet been reread against the new exclusion

One flagged code got anchors and no rewrite. The other got a real exclusion rule, and still owes a rescore on 27 rows.

What's in the pack

01

Reliability Procedure

Fixes the coders, the double-coding sample size, the statistic and the flag threshold before anyone codes anything, so the threshold cannot be picked after seeing which codes look bad.

02

Agreement by Code sheet

Every code's own raw agreement, kappa, interpretation band and flag, scored separately against the same excerpts rather than pooled into one blended figure.

03

Disagreement Register sheet

Every disputed extract behind a flagged code, quoted verbatim from the same cleaned transcripts the codebook was built from, with why the two coders' calls split.

04

Codebook Refinement Log sheet

What changed in the codebook and why, or the decision that nothing needed to, naming the exact register rows behind each call.

05

Disagreement Resolution Note

The narrative read of each flagged code's register, naming which disagreements shared one pattern and which were independent calls on a rare topic. Reporting checklists ask directly whether reliability was assessed, not only whether a number exists.

06

A flagged code gets read before it gets rewritten

The standing rule every prompt reads first. A code's definition changes only after its disagreements are read, never on the strength of a low score alone.

How to use it

  1. 1

    Open in River, or download it

    Open the pack in River and the agent scores the codes with you, or download the two Word documents and three CSV sheets instantly, filled in with the worked example.

  2. 2

    Send both coders' independent output

    The codebook and each coder's calls on the same double-coded excerpts, coded without either one seeing the other's calls first.

  3. 3

    Read the register on whatever gets flagged

    Agreement by Code scores every code and flags whichever fall under the threshold. Disagreement Register holds their disputed extracts, quoted verbatim, before any definition changes.

  4. 4

    Log the change, or the decision not to

    Codebook Refinement Log records what happened and why. Once reliability is settled, theme development starts from codes that held up under a second coder.

Frequently asked questions

Is this template free?

Yes. Download the whole pack as Word documents and CSV sheets with no signup and no credit card. Edit with AI is a separate, optional path that has the agent score the codes and build the register with you. The template library holds the rest of the packs.

My QDA software already calculates a reliability coefficient. Why use this?

NVivo and ATLAS.ti will compute a coefficient for you, usually pooled across the whole coding frame in one figure. Neither flags a specific code as needing a second look, and neither holds the disputed extracts behind a low score. This pack scores every code separately and routes only the flagged ones into a register before anyone touches a definition.

What format are the downloaded files?

CSV for the three sheets and Word documents for Reliability Procedure and Disagreement Resolution Note, in one zip. They open in Excel, Numbers, Sheets, Word, Pages and Google Docs with nothing to convert. Inside River the same content opens as live Docs and Sheets.

How is this different from a plain kappa calculator?

A calculator returns one number for one 2x2 table you already built by hand. This pack scores every code in a codebook against the same double-coded excerpts, flags whichever fall short of a threshold set before scoring, and pulls each flagged code's actual disputes into a register rather than stopping at the coefficient.

What does 'Edit with AI' actually do?

It creates a free River account and installs this exact pack as a private workspace. Send the codebook and each coder's independent output, and the agent scores every code, builds the register for whichever fall short, and drafts the resolution note from what you decide.

Do I need exactly two coders?

The worked example and the standing default, Cohen's kappa, are built for two. With three or more coders, use Fleiss' kappa or Krippendorff's alpha instead. River still scores each code separately against the same threshold and still builds the register the same way; only the formula behind the number changes.

How much of my data actually needs to be double-coded?

Something in the range of 10 to 25 percent of the full codable set is typically enough for a trustworthy estimate, drawn by a stated, repeatable rule rather than either coder picking which excerpts look easy. Run a small pilot first, as little as one interview, before committing to the full sample.

Find out which of your codes actually needs a second look

Send the codebook and each coder's independent output. The register comes back naming which flagged code is a paradox and which is a real gap.

Edit with AI