Grade Analysis and Moderation Template
Two documents and three sheets that read every section's grade against a score no marker's judgment can touch.
Free download · No account needed
Distribution by Marker and Section
Every section's grade read against a score no marker could touch
One row per section. The residual is graded performance minus what the anchor predicts, reported in standard deviations, points and this course's own grading bands.
| Marker | Section n | Part A mean | Part A SD | Part B mean | Part B SD | Residual (points) | Residual (bands) | Flag |
|---|---|---|---|---|---|---|---|---|
Two faults, two thresholds, set before any gradebook is read
A residual of 0.25 bands or more in either direction is a level flag. A graded-component standard deviation under 65 percent of the cohort median is a spread flag, a different problem a numeric correction cannot fix.
A section's average grade is a fact about its students until it is read against a score no marker's judgment touches. On the worked cohort in this pack, an economics course's final exam, 357 students sit across 6 sections, each assessment split between a multiple-choice half and an essay half. One section grades 0.37 of a band harder than its own students' multiple-choice performance predicts, and a second shows a different fault: essay scores landing within a couple of points of each other regardless of quality.
Ohio State's guidance for large-enrollment courses names the fix directly: compare a grader's subjective scores against a more objective measure from the same assessment, such as the multiple-choice half, rather than comparing raw section averages to each other. That comparison is the residual this pack computes for every section, in standard deviations, in points, and in the course's own grading bands, so a department reads a number it already knows how to act on.
Built for department chairs and course coordinators moderating a multi-section assessment marked by more than one person, and for anyone who has watched two sections drift apart and could not tell whether the students or the marking caused it. Before trusting a flag, item analysis review is where to check that the anchor's own questions are behaving. Where the shared rubric itself is the harder problem, the rubric pack is where it gets rewritten.
What's in the pack
Moderation Procedure
The full method: why a raw average is not evidence, how the residual is computed and reported three ways, and the level and spread thresholds stated before any gradebook is read.
Distribution by Marker and Section
One row per section: both halves' means and standard deviations, the residual in standard deviations, points and bands, and which threshold a section crossed, set against the whole cohort's own figures.
Outlier Register
Every flagged section with its signal, its evidence and a recommended action, separating a level fault a blind remark can confirm from a spread fault that a numeric correction cannot touch at all.
Moderation Adjustment Log
The blind remark sample against the original scores, the direction and size of the shift, and the section-wide correction it justified, so the correction traces back to a second marker rather than a formula alone.
Findings Note
The department-facing writeup: which sections were flagged, what the blind remark found, and what was and was not changed, with the arithmetic left in rather than summarized away.
The anchor and the rubric
This pack tests whether a section was marked consistently, not whether the rubric or the anchor's own questions are sound. The feedback efficiency pack is where the comments themselves get faster to write.
How it works
- 1
Send the gradebook
Every student's marker or section, their graded score, and a score from a component nobody's judgment touches, from the same assessment.
- 2
Read the distribution
Every section's residual against the anchor, reported in standard deviations, points and this course's own grading bands, next to the cohort's own figures.
- 3
Check the flags, not just the averages
A level flag at 0.25 bands or more, a spread flag under 65 percent of the median standard deviation, each needing a different fix.
- 4
Remark blind before correcting anything
A second marker rescores a sample without seeing the original marks, and only a consistent result becomes a section-wide correction.
Frequently asked questions
Is this template free?
Yes. The two documents and three sheets download as Word and CSV files with no signup and no card. "Edit with AI" is the optional path for departments that want the residuals computed from their own gradebook export. The rest of the library is at the template library.
What format are the downloaded files?
Word documents for the procedure and the findings note, plus CSV for the three sheets, zipped into one file. They open in Word, Pages or Google Docs and Excel, Numbers or Sheets, with nothing to convert before you can read them.
What counts as an anchor, and what if we don't have one?
Any component of the same assessment scored the same way regardless of section: multiple choice against a single key, short answer against a fixed list, or a machine-graded portion. Without one, there is nothing to read the graded scores against, and the method does not run.
Why not just compare each section's raw average?
Because a marker who draws the stronger half of the cohort shows a higher average than a harsher marker who draws the weaker half, and a raw comparison cannot separate those two cases. Reading each section against its own anchor performance is what makes the comparison marker against ability rather than marker against marker.
Does a flagged section get its grades changed automatically?
No. A level flag is confirmed with a blind remark, a sample a second marker scores without seeing the original marks, before a single grade changes. A spread flag is routed to a norming session instead, since there is no single direction to shift a compressed section by.
Can this tell a department a marker should stop marking?
No, and it does not adjudicate one student's appeal either. An appeal is a claim about one paper, decided by rereading that paper, not by a section-wide statistic. What it produces is the distribution, the flags and the confirmed findings a department reasons from.
How is this different from an item analysis?
An item analysis asks whether one question on the anchor itself discriminates between strong and weak students, which is worth checking before trusting the anchor. Item analysis review is the right tool for that. This asks whether a marker's grading tracks the anchor once it is trusted.
Find the section that's grading off the line
Send the gradebook export with a marker-independent score alongside the graded one. The first pass returns every section's residual and which ones crossed a threshold.
Moderate a real cohort