Education & TrainingFree
Item Analysis Interpretation for Exams
Send the item analysis export and get every item sorted into keep, revise, rekey or remove, with the specific defect named.
River's item analysis review takes the report your testing service or learning platform produced and does the part it leaves to you. Every item gets a verdict rather than a statistic: keep, revise, rekey or remove. The verdict comes from crossing three things the report prints separately, which are percent correct, the discrimination column, and how each option's takers scored overall. You get a sheet with one row per item and a document naming the defect in each flagged one.
Unlike the report itself, which sorts by difficulty and stops, this reads the two directions a difficulty sort cannot. Downward, it separates an item that is hard from an item that is broken, using discrimination rather than percent correct. Upward, it finds the items with a comfortable percent correct and no discrimination at all, which never appear near the top of any list you were given and are the ones measuring nothing. Both readings come out of the same three columns the report already prints.
It is built for the instructor holding a report they did not ask for, and for the assessment coordinator who has to say something defensible about an exam before grades post. Use it after a midterm when a question drew complaints, before reusing a bank from a previous term, or when a departmental review asks which items were retired and why. The rubric pack covers the graded work that is not multiple choice.
Why percent correct cannot tell you an item is broken
The single most common mistake in reading one of these reports is treating the difficulty column as a quality column. It is not one. An item everybody misses can be an excellent item on a genuinely hard concept, and an item two thirds of the class gets right can be measuring nothing at all. What separates them is discrimination: whether the students who did well on the exam did well on this item. UT Austin's testing service reads 0.30 and above as a good item and says to remove anything at or below zero.
Ravensdale Community College, BIOL 211, a 48 item midterm sat by 186 students. Ten items fall below 50% correct, and six of them discriminate at 0.35 or better, so they are hard and worth keeping. Four are broken. Q39 is the clearest: 30% chose the recorded key D, which discriminates at minus 0.33, while 48% chose option A, taken by 74% of the top 50 scorers and 12% of the bottom 50. A discriminates at plus 0.50. The key is wrong.
Then read the list the other way. Five items at Ravensdale have a comfortable percent correct and discrimination under 0.10, three of them below zero, including one answered correctly by 80% of the class. The 27 items in the same difficulty range that do work average 0.52. None of the five appears near the top of a report sorted by difficulty, so nobody looks at them. Sort by discrimination instead and they arrive first, which is the whole trick.
How it works
Paste the report
The item analysis export as it came out, plus what the exam was and how many students sat it.
Name the columns
Percent correct alone, or discrimination too, or option counts, or option counts split by score group.
A verdict per item
Hard items kept, broken ones named by defect, and the items a difficulty sort never surfaced.
Rewrite the flagged stems
Ask for a rewritten stem, a replacement distractor, or the score effect of rekeying one item.
What you get
- One row per item with a verdict of keep, revise, rekey or remove and the evidence behind it
- Hard items separated from broken ones by discrimination, so a good difficult question survives review
- Every option scored for share and direction, naming which distractor the prepared students chose
- Miskeyed items caught where a distractor out-discriminates the answer key the report used
- Non-functioning distractors counted, with the guessing floor your option count actually delivers
- The rekey consequence in grades: how many raw scores move and how many letters change
Common questions
My report only gives percent correct. Is this any use?
Partly, and it will tell you which part. Without a discrimination column no item can be separated into hard or broken, so you get the difficulty distribution, the items outside the usual bands, and a list of which column your platform can export to make the rest possible. Most of them can export it.
What counts as a bad discrimination value?
The conventional reading treats 0.40 and above as very good, 0.30 to 0.39 as good and 0.19 or below as poor. Below zero the item is backwards, because the weaker students outperformed the stronger ones on it. Sample size matters: under 30 examinees these figures are unstable and get flagged rather than reported.
How do you know an item is miskeyed rather than just hard?
By direction. On a hard item the key still discriminates positively, because prepared students find it. On a miskeyed item a distractor discriminates more positively than the key, so the top scorers are choosing that option and the weaker ones are choosing what the answer sheet says. That reversal has one cause.
Does a distractor nobody picks actually matter?
It matters to the guessing floor. A distractor is doing work only if it is chosen by at least 5% of examinees and pulls the weaker students. At Ravensdale 27 of 144 distractors failed that test, which moved the chance score from 12 marks to 16.9 out of 48.
Should I throw out an item everyone got wrong?
Not on percent correct alone. Six of the ten hardest items at Ravensdale discriminated at 0.35 or better, which means they were separating students who knew the material from students who did not. Removing those makes the exam less informative. Remove the four that discriminated near or below zero.
Can it tell me whether to drop an item after grades are in?
It gives you the arithmetic for that decision and leaves the decision with you. For a miskeyed item it computes how many raw scores move and how many letter grades change in each direction. At Ravensdale, rekeying two items moved 29 letter grades, 26 of them up.
Where does the rest of the exam work happen?
This tool reads a report about an exam that already ran. Writing the outcomes the items answer to is the learning outcomes pack, and the alignment between outcomes and assessments comes from the course design pack. Organising the bank the items came from, and working out whether it can fill the next blueprint, is the question bank pack.
Item Analysis Interpretation for Exams
Fill in the form and your workspace opens with the work already underway.