Intercoder Reliability Template for Research
Two documents and three sheets that score every code separately, flag whichever fall short, and hold each flagged code's disputed extracts in a register.
Free download · No account needed
Agreement by Code
Five codes, scored separately against the same 120 excerpts
| Code | Raw agreement | Kappa | Band | Flag |
|---|---|---|---|---|
| Described mentor check-ins as helpful | 90.8% | 0.815 | Almost perfect | No |
| Requested more classroom-management support | 90.0% | 0.100 | Slight | YES |
| Described the program as effective overall | 70.8% | 0.391 | Fair | YES |
| Named a specific curriculum resource received | 97.5% | 0.908 | Almost perfect | No |
| Named a specific administrative barrier to implementation | 95.0% | 0.640 | Substantial | No |
Two codes sit within a point of each other on raw agreement, 90.8% and 90.0%. Their kappas are 0.815 and 0.100, eighteen bands apart.
A dual-coded qualitative study usually reports one number: a kappa next to the codebook, footnoted that disagreements were discussed and resolved. Pooling every code into that figure hides which code it punishes. Cohen's kappa corrects for the agreement two coders would reach by guessing, and that correction has a documented failure: a rare or skewed code can return a kappa near zero from a handful of ordinary calls. A published appraisal of 57 clinical trials found exactly this: 84.2 percent raw agreement produced a kappa of only 0.042, because one category covered nearly every trial.
The pack is two documents and three sheets. Reliability Procedure fixes the coders, the sample size, the statistic and the flag threshold before anyone codes anything, so a threshold never gets picked after seeing which codes look bad. Agreement by Code scores every code separately against the excerpts and flags whichever fall under that line, read against the standard bands rather than in isolation. Disagreement Register holds the disputed extracts behind every flagged code, quoted verbatim with why the coders split. Codebook Refinement Log records what changed, or that nothing needed to.
Reading the register is what tells a paradox apart from a real gap. In the worked example, one flagged code's twelve disagreements share no pattern, so two anchor examples get added and the definition stays put. A second flagged code's disagreements repeat one boundary problem six times in an eight-row sample, so an exclusion rule gets written instead, and Disagreement Resolution Note records which happened and why. Open the pack in River and the agent scores the codes with you, or download the sheets and read your own codebook's disagreements the same way.
What's in the pack
Reliability Procedure
Fixes the coders, the double-coding sample size, the statistic and the flag threshold before anyone codes anything, so the threshold cannot be picked after seeing which codes look bad.
Agreement by Code sheet
Every code's own raw agreement, kappa, interpretation band and flag, scored separately against the same excerpts rather than pooled into one blended figure.
Disagreement Register sheet
Every disputed extract behind a flagged code, quoted verbatim from the same cleaned transcripts the codebook was built from, with why the two coders' calls split.
Codebook Refinement Log sheet
What changed in the codebook and why, or the decision that nothing needed to, naming the exact register rows behind each call.
Disagreement Resolution Note
The narrative read of each flagged code's register, naming which disagreements shared one pattern and which were independent calls on a rare topic. Reporting checklists ask directly whether reliability was assessed, not only whether a number exists.
A flagged code gets read before it gets rewritten
The standing rule every prompt reads first. A code's definition changes only after its disagreements are read, never on the strength of a low score alone.
How to use it
- 1
Open in River, or download it
Open the pack in River and the agent scores the codes with you, or download the two Word documents and three CSV sheets instantly, filled in with the worked example.
- 2
Send both coders' independent output
The codebook and each coder's calls on the same double-coded excerpts, coded without either one seeing the other's calls first.
- 3
Read the register on whatever gets flagged
Agreement by Code scores every code and flags whichever fall under the threshold. Disagreement Register holds their disputed extracts, quoted verbatim, before any definition changes.
- 4
Log the change, or the decision not to
Codebook Refinement Log records what happened and why. Once reliability is settled, theme development starts from codes that held up under a second coder.
Frequently asked questions
Is this template free?
Yes. Download the whole pack as Word documents and CSV sheets with no signup and no credit card. Edit with AI is a separate, optional path that has the agent score the codes and build the register with you. The template library holds the rest of the packs.
My QDA software already calculates a reliability coefficient. Why use this?
NVivo and ATLAS.ti will compute a coefficient for you, usually pooled across the whole coding frame in one figure. Neither flags a specific code as needing a second look, and neither holds the disputed extracts behind a low score. This pack scores every code separately and routes only the flagged ones into a register before anyone touches a definition.
What format are the downloaded files?
CSV for the three sheets and Word documents for Reliability Procedure and Disagreement Resolution Note, in one zip. They open in Excel, Numbers, Sheets, Word, Pages and Google Docs with nothing to convert. Inside River the same content opens as live Docs and Sheets.
How is this different from a plain kappa calculator?
A calculator returns one number for one 2x2 table you already built by hand. This pack scores every code in a codebook against the same double-coded excerpts, flags whichever fall short of a threshold set before scoring, and pulls each flagged code's actual disputes into a register rather than stopping at the coefficient.
What does 'Edit with AI' actually do?
It creates a free River account and installs this exact pack as a private workspace. Send the codebook and each coder's independent output, and the agent scores every code, builds the register for whichever fall short, and drafts the resolution note from what you decide.
Do I need exactly two coders?
The worked example and the standing default, Cohen's kappa, are built for two. With three or more coders, use Fleiss' kappa or Krippendorff's alpha instead. River still scores each code separately against the same threshold and still builds the register the same way; only the formula behind the number changes.
How much of my data actually needs to be double-coded?
Something in the range of 10 to 25 percent of the full codable set is typically enough for a trustworthy estimate, drawn by a stated, repeatable rule rather than either coder picking which excerpts look easy. Run a small pilot first, as little as one interview, before committing to the full sample.
Find out which of your codes actually needs a second look
Send the codebook and each coder's independent output. The register comes back naming which flagged code is a paradox and which is a real gap.
Edit with AI