River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Critical Appraisal Tool Template

Three documents and three sheets that keep two reviewers' independent ratings and evidence apart until a disagreement is resolved, never averaged.

Free download  ·  No account needed

Appraisal by Study and Domain

[Study], [Domain]

Reviewer AReviewer BEvidence quotedConsensus
[Low / Some concerns / High][Low / Some concerns / High][The exact passage each reviewer read][Filled in only once both ratings are in]

One row per domain per study. Each reviewer appraises before seeing the other's column, so the agreement check that runs afterward means something.

Most critical appraisal checklists, CASP, JBI, AMSTAR-2, RoB 2 itself, hand you the same thing: five or so domains, a three-point scale, and a place to write the rating. What happens when two reviewers rate a study differently is left to the review team, and Cochrane's own guidance for RoB 2 says the individual ratings from each reviewer do not need to be kept, only the agreed, consolidated ones. This pack keeps them anyway, in the same three sheets and three documents, on purpose.

The pack ships filled in for a review of structured mentoring programmes and new-hire turnover: nine studies, two independent reviewers, forty-five domain ratings. Agreement at the domain level is 82.2 per cent, which alone would read as solid. Agreement on the overall verdict, the number a reader actually sees, is only 55.6 per cent, because the overall rating is the least favourable assessment across the five domains rather than an average. One domain, Confounding, carries five of the eight disagreements alone.

This space rates studies a review has already assembled, working domain by domain rather than by an overall impression, then running the agreement check before touching the risk-of-bias summary. It does not run the search, which the search strategy pack documents while it is still running, or the screening, which the PRISMA screening template turns into a reproducible flow. It also does not pool effect sizes into a meta-analysis; a synthesis reads these domain ratings to weight studies rather than average them as equal.

Every sheet in the pack

Appraisal by Study and Domain, Reviewer Agreement Check, and Risk of Bias Summary.

Appraisal by Study and Domain

4 of the 8 domain-level disagreements in the worked example, of 45 ratings total.

StudyDomainRev. ARev. BConsensus
Larkspur 2024MeasurementSome concernsHighHigh
Okonkwo 2023ConfoundingHighSome concernsHigh
Whitfield 2021ConfoundingLowSome concernsSome concerns
Halvorsen 2023ReportingLowSome concernsSome concerns

Whitfield's disagreement never changes its overall verdict, already High from Selection. Halvorsen's does not either, but it adds a second driver domain to a verdict that already matched.

Reviewer Agreement Check

Domain by domain first, then the overall verdict, on the same 9 studies.

DomainAgreementsRate
Selection9 of 9100.0%
Measurement8 of 988.9%
Confounding4 of 944.4%
Attrition8 of 988.9%
Reporting8 of 988.9%
All domains37 of 4582.2%
Overall verdict5 of 955.6%

Confounding alone carries 5 of the 8 domain disagreements. The overall-verdict rate is lower than the domain rate because the verdict is a worst-of-five, not an average: one bad domain is enough to flip it.

Risk of Bias Summary

Consensus only, after every flagged disagreement was resolved.

StudySelect.Meas.Confound.Attrit.Report.Overall
Okafor 2023LowLowLowLowLowLow
Bramwell 2022SomeLowLowSomeLowSome
Larkspur 2024LowHighSomeLowLowHigh
Whitfield 2021HighSomeSomeSomeLowHigh
Okonkwo 2023LowLowHighLowSomeHigh
Delacroix 2022SomeLowHighSomeLowHigh
Marchetti 2024LowLowLowLowLowLow
Suzuki 2021SomeLowSomeHighSomeHigh
Halvorsen 2023LowSomeLowLowSomeSome

2 Low, 2 Some concerns, 5 High overall. Confounding drives 2 of the 5 High studies outright; Selection, Measurement and Attrition each drive one of the rest.

What's in the pack

01

Appraisal by Study and Domain sheet

Both reviewers' ratings and quoted evidence, side by side, never overwritten by the consensus. The resolution note lives in the same row as the disagreement it explains.

02

Reviewer Agreement Check sheet

Agreement computed domain by domain before it is computed on the overall verdict, with every disagreeing study named rather than folded into one blended rate.

03

Risk of Bias Summary sheet

The consensus table, populated only once every flagged disagreement has a resolution note, so a rating that is still contested cannot quietly reach the summary.

04

Risk of Bias Plot

A traffic-light table by study and domain plus the per-domain rating distribution, regenerated from the summary sheet rather than hand-drawn.

05

Appraisal Criteria

Five domains, Selection through Reporting, each with three anchored examples drawn from named studies in your own review rather than generic placeholders.

06

Reviewer Guidance

Rules for a dual-review team: appraise before comparing, quote before rating, and resolve a disagreement by rereading the study together rather than splitting the difference.

07

Two ratings and the evidence, never one rating and a checkbox

The standing space rule every prompt reads first, which is why the sheet keeps both reviewers' rows instead of collapsing to the agreed rating on entry.

How to use it

  1. 1

    Send the studies

    The full text where possible, since Confounding and Attrition ratings usually depend on a passage in the methods or results, not the abstract.

  2. 2

    Appraise independently, domain by domain

    Each reviewer rates Selection through Reporting on their own copy, quoting the passage before choosing the rating, without seeing the other reviewer's row.

  3. 3

    Run the agreement check before the summary

    Domain by domain first, then the overall verdict, so a domain carrying most of the disagreement is visible before it disappears into a consensus number.

  4. 4

    Resolve by rereading, then generate the plot

    Every disagreement gets a resolution note from rereading the study together, or a third reviewer's independent read, before Risk of Bias Summary and the plot regenerate.

Frequently asked questions

Is this template free?

Yes. Download the whole pack as Word documents and CSV sheets with no signup and no credit card. Edit with AI is a separate, optional path that has the agent run the appraisal with you, keeping both reviewers' ratings apart until a disagreement is resolved. The template library holds the rest of the packs.

What format are the downloaded files?

Word documents for the three write-ups and CSV for the three sheets, in one zip. They open natively in Word, Pages, Google Docs, Excel, Numbers and Sheets with nothing to convert. Inside River the same content opens as live Docs and Sheets, with both reviewers editing the same appraisal sheet at once.

Is this the same as the RoB 2 tool or a JBI checklist?

No. RoB 2, JBI's checklists, CASP and AMSTAR-2 define the domains and the rating scale; this pack does not replace any of them and its Appraisal Criteria is written to be adapted to whichever one your review already uses. What it adds is the sheet that keeps two reviewers' ratings apart, which none of those checklists' own templates provide.

Does River appraise the studies for me?

No. Rating a study's risk of bias is a judgement call that has to be made by someone who read it, and the value of dual review is that two people made that judgement independently. River structures the process: it holds the criteria, keeps both reviewers' evidence and ratings separate, runs the agreement check, and writes up the resolution once you decide it.

Do we need two reviewers?

No. A single-reviewer appraisal is common for smaller or lower-stakes reviews, and this pack works that way too: Reviewer Agreement Check simply stays empty, and Risk of Bias Summary is populated straight from the one reviewer's ratings. Add a second reviewer later and the agreement check activates on whatever has already been rated.

What happens if the two reviewers still disagree after rereading the study?

Bring in a third reviewer who rates the disputed domain independently, from the source study, without seeing either of the first two ratings, then all three discuss. That is standard dual-review practice for exactly this case: a genuine, defensible difference in how two trained readers interpret the same passage, which a third independent read settles better than averaging two numbers.

Keep both reviewers' ratings, not just the agreed one

Send the studies you are appraising and whether one reviewer or two will rate them. The evidence gets quoted before the rating, and every disagreement gets a resolution note.

Edit with AI