River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Grade Analysis and Moderation Template

Two documents and three sheets that read every section's grade against a score no marker's judgment can touch.

Free download  ·  No account needed

Distribution by Marker and Section

Every section's grade read against a score no marker could touch

One row per section. The residual is graded performance minus what the anchor predicts, reported in standard deviations, points and this course's own grading bands.

MarkerSection nPart A meanPart A SDPart B meanPart B SDResidual (points)Residual (bands)Flag
         
         

Two faults, two thresholds, set before any gradebook is read

A residual of 0.25 bands or more in either direction is a level flag. A graded-component standard deviation under 65 percent of the cohort median is a spread flag, a different problem a numeric correction cannot fix.

A section's average grade is a fact about its students until it is read against a score no marker's judgment touches. On the worked cohort in this pack, an economics course's final exam, 357 students sit across 6 sections, each assessment split between a multiple-choice half and an essay half. One section grades 0.37 of a band harder than its own students' multiple-choice performance predicts, and a second shows a different fault: essay scores landing within a couple of points of each other regardless of quality.

Ohio State's guidance for large-enrollment courses names the fix directly: compare a grader's subjective scores against a more objective measure from the same assessment, such as the multiple-choice half, rather than comparing raw section averages to each other. That comparison is the residual this pack computes for every section, in standard deviations, in points, and in the course's own grading bands, so a department reads a number it already knows how to act on.

Built for department chairs and course coordinators moderating a multi-section assessment marked by more than one person, and for anyone who has watched two sections drift apart and could not tell whether the students or the marking caused it. Before trusting a flag, item analysis review is where to check that the anchor's own questions are behaving. Where the shared rubric itself is the harder problem, the rubric pack is where it gets rewritten.

One section grading 0.37 of a band harder than its own students earned

The per-section distribution against the anchor, the two flagged sections with their evidence, and the blind remark that confirmed one of them.

Distribution by Marker and Section

ECON 301 final exam, 357 students across 6 sections. Part A is 60 points of multiple choice on one key; Part B is a 40-point essay marked by each section's own TA. One band is 3 points, this course's convention.

MarkernPart A meanPart B meanPart B SDResidual (pts)Residual (bands)Flag
Ferris5744.629.44.6+0.19+0.06
Delacroix6145.329.95.1+0.23+0.08
Nkemelu5844.928.34.4−1.11−0.37Level (harsh)
Voss6245.730.35.4+0.36+0.12
Iyer5944.229.04.9+0.06+0.02
Castellano6045.029.72.6+0.23+0.08Spread (compressed)
COHORT35744.9629.454.640.000.00Reference row

Castellano's residual is ordinary, +0.08 bands, so level is not the fault. The tell is Part B SD: 2.6 against a cohort median of 4.75, below the 3.09 floor (65 percent of median), meaning strong and weak papers land within a few points of each other.

Outlier Register

The two sections that crossed a threshold, with the evidence and the action, before either marker knows they were flagged.

RefMarkerSignalEvidenceSeverityAction
O-1NkemeluLevel, graded below what Part A predicts−0.37 bands (−1.11 pts, −0.24 sd)HighDraw a blind remark sample before correcting anything
O-2CastellanoSpread, graded scores compressedSD 2.6 vs. median 4.75, floor 3.09MediumFlag for the next calibration session

Nothing here is a correction yet. A level flag has to survive a blind remark first, and a spread flag never gets a numeric correction at all, because there is no single direction to shift a compressed section by.

Moderation Adjustment Log

Ten of Nkemelu's 58 papers, the larger of ten or 15 percent, rescored by a second marker blind to the original grade.

StudentOriginal (/40)Remark (/40)Difference
S-1181819+1
S-1222224+2
S-1272425+1
S-1312627+1
S-1342830+2
S-1392930+1
S-1423132+1
S-14533330
S-1493536+1
S-1533739+2
9 up / 1 unchanged / 0 down  +1.2 mean

Mean uplift 1.2 points against the residual's own prediction of 1.11, close enough and consistent enough to apply a flat +1.2 correction to all 58 graded scores rather than remarking the remaining 48 (2.5 hours sampled against an estimated 14.5 hours full).

What's in the pack

01

Moderation Procedure

The full method: why a raw average is not evidence, how the residual is computed and reported three ways, and the level and spread thresholds stated before any gradebook is read.

02

Distribution by Marker and Section

One row per section: both halves' means and standard deviations, the residual in standard deviations, points and bands, and which threshold a section crossed, set against the whole cohort's own figures.

03

Outlier Register

Every flagged section with its signal, its evidence and a recommended action, separating a level fault a blind remark can confirm from a spread fault that a numeric correction cannot touch at all.

04

Moderation Adjustment Log

The blind remark sample against the original scores, the direction and size of the shift, and the section-wide correction it justified, so the correction traces back to a second marker rather than a formula alone.

05

Findings Note

The department-facing writeup: which sections were flagged, what the blind remark found, and what was and was not changed, with the arithmetic left in rather than summarized away.

06

The anchor and the rubric

This pack tests whether a section was marked consistently, not whether the rubric or the anchor's own questions are sound. The feedback efficiency pack is where the comments themselves get faster to write.

How it works

  1. 1

    Send the gradebook

    Every student's marker or section, their graded score, and a score from a component nobody's judgment touches, from the same assessment.

  2. 2

    Read the distribution

    Every section's residual against the anchor, reported in standard deviations, points and this course's own grading bands, next to the cohort's own figures.

  3. 3

    Check the flags, not just the averages

    A level flag at 0.25 bands or more, a spread flag under 65 percent of the median standard deviation, each needing a different fix.

  4. 4

    Remark blind before correcting anything

    A second marker rescores a sample without seeing the original marks, and only a consistent result becomes a section-wide correction.

Frequently asked questions

Is this template free?

Yes. The two documents and three sheets download as Word and CSV files with no signup and no card. "Edit with AI" is the optional path for departments that want the residuals computed from their own gradebook export. The rest of the library is at the template library.

What format are the downloaded files?

Word documents for the procedure and the findings note, plus CSV for the three sheets, zipped into one file. They open in Word, Pages or Google Docs and Excel, Numbers or Sheets, with nothing to convert before you can read them.

What counts as an anchor, and what if we don't have one?

Any component of the same assessment scored the same way regardless of section: multiple choice against a single key, short answer against a fixed list, or a machine-graded portion. Without one, there is nothing to read the graded scores against, and the method does not run.

Why not just compare each section's raw average?

Because a marker who draws the stronger half of the cohort shows a higher average than a harsher marker who draws the weaker half, and a raw comparison cannot separate those two cases. Reading each section against its own anchor performance is what makes the comparison marker against ability rather than marker against marker.

Does a flagged section get its grades changed automatically?

No. A level flag is confirmed with a blind remark, a sample a second marker scores without seeing the original marks, before a single grade changes. A spread flag is routed to a norming session instead, since there is no single direction to shift a compressed section by.

Can this tell a department a marker should stop marking?

No, and it does not adjudicate one student's appeal either. An appeal is a claim about one paper, decided by rereading that paper, not by a section-wide statistic. What it produces is the distribution, the flags and the confirmed findings a department reasons from.

How is this different from an item analysis?

An item analysis asks whether one question on the anchor itself discriminates between strong and weak students, which is worth checking before trusting the anchor. Item analysis review is the right tool for that. This asks whether a marker's grading tracks the anchor once it is trusted.

Find the section that's grading off the line

Send the gradebook export with a marker-independent score alongside the graded one. The first pass returns every section's residual and which ones crossed a threshold.

Moderate a real cohort