Scale Validation Procedure Template
Two documents and four sheets that connect each item's fate, kept, reworded or dropped, to the exact statistic that decided it
Free download · No account needed
Item Revision Log
Change Readiness Scale, Highmark Logistics pilot
| Item | Statistic that fired | Decision |
|---|---|---|
| C4 | Item-total r=.20; alpha rises to .84 if dropped; loading .34, communality .13 | Drop |
| E5 | Item-total r=.44 and alpha-if-excluded=.82 both pass; loading .54 lands on the wrong factor | Reword, hold for revalidation |
| B4 | Item-total r=.59; alpha falls, not rises, if excluded | Retain |
C4 fails all three sheets. E5 passes two and fails the third, the case this log exists to catch. B4 is reviewed and stays.
A scale-validation tutorial states the actual decision rule in one row of one table. A corrected item-total correlation under .30, or a rising Cronbach's alpha, is reason enough to flag an item, a rule a 2018 protocol in Health Psychology and Behavioral Medicine states inside one study's results narrative. A 2024 best-practice review adds a second test: whether an item's factor loading lands on a different construct once every item runs against every factor together. Neither names a place to log an item that clears the first test and fails the second.
This space fixes those checks as three standing thresholds before a single pilot number is read, then logs every item's fate to the exact statistic that decided it. Item-total and reliability are both computed against an item's own subscale in isolation, so an item can clear both while still measuring the wrong construct; only the factor-structure check tests it against everything else in the pool at once. It does not draft the item pool itself. A dedicated instrument-design pack does that, tracing every item to a construct before a single response exists.
In the worked example, Highmark Logistics pilots an 18-item Change Readiness Scale on 238 employees ahead of a warehouse system changeover. C4 fails every sheet: an item-total correlation of .20, an alpha that rises to .84 without it, and a factor loading of just .34 on its own construct. E5 clears item-total and reliability cleanly but loads .54 on the wrong factor against .42 on its own, the pattern only the third check catches, so it is reworded and held for revalidation rather than treated as finished.
What's in the pack
Method Note
The three decision thresholds, item-total, reliability and factor structure, fixed and named before a single item statistic from the actual pilot is computed.
Item Statistics sheet
Every item's mean, standard deviation and corrected item-total correlation against its own subscale, with anything under .30 flagged in the same row.
Reliability by Subscale sheet
Cronbach's alpha as drafted for each subscale, the item whose removal would raise it most, and the revised alpha once that item is actually dropped or kept.
Factor Structure sheet
Every item's loading on every factor at once, not just its own, with communality and a flag for a loading that lands outside its home factor by more than a minor gap.
Item Revision Log sheet
Which threshold fired for every flagged item, the exact number that fired it, and the decision, drop, reword or retain, with one summary row for everything that cleared all three.
Validation Narrative
The pattern read in prose once every item's fate is settled: what was dropped, what was reworded and why, and what one pilot does not establish on its own.
An item's fate traces to every sheet that tested it, not just the one that passed
The standing space rule every prompt reads first, including why two passing numbers do not rule out a factor-structure problem.
How to use it
- 1
Open in River, or download it
Open the pack in River and the agent fixes the thresholds and runs the checks against your own pilot data, or download the two documents and four CSV sheets instantly, filled in with the worked example.
- 2
Send the item pool and the pilot data
The drafted items grouped by the construct or subscale each one was written for, and the pilot's response data in rows and columns, however it was collected.
- 3
Thresholds get fixed before a number is read
Item-total, reliability and factor-structure thresholds go into Method Note before a single statistic from the actual pilot is computed, so no threshold gets picked after seeing which items would fail it.
- 4
Every item's fate gets logged, then read as one pattern
Item Revision Log states which threshold fired and what the number was for every flagged item, and Validation Narrative reads what the full pattern does and does not establish.
Frequently asked questions
Is this template free?
Yes. Download the whole pack as a Word document and CSV sheets with no signup and no credit card, filled in for the Highmark Logistics worked example. Edit with AI is a separate, optional path that fixes the thresholds and runs the checks against your own pilot data. The template library holds the rest of the packs.
What format are the downloaded files?
Word documents for Method Note and Validation Narrative, and CSV for the four sheets, all in one zip. They open in Word, Pages, Google Docs, Excel, Numbers and Sheets with nothing to convert. Inside River the same content opens as live Docs and Sheets.
Does this draft the survey items, or decide what the scale measures?
No. It assumes the item pool and the pilot data already exist and decides which items keep their place, get reworded, or get dropped. A dedicated instrument-design pack traces every item to a construct and a planned analysis before a single response is collected, which is a separate, earlier step.
How small can the pilot sample be?
The sources behind this template's thresholds treat a participant-to-item ratio between 5 to 1 and 10 to 1 as commonly followed for a factor analysis. The worked example pilots 238 respondents against 18 items, a 13 to 1 ratio. Below that range, item statistics and reliability still run, but the factor-structure check becomes the least trustworthy of the three.
What happens to an item that clears two checks and fails the third?
That is the case this template is built to catch. Item-total correlation and reliability are both computed against an item's own subscale in isolation, so an item can pass both while its factor loading lands on a different construct once tested against the whole pool. The item gets reworded and held for revalidation, not waved through.
Does it run the statistics for me, or do I need the pilot data already collected?
It needs your pilot's response data already in rows and columns, one column per item. If the export still carries a survey platform's raw formatting, the export preparation tool handles the platform quirks first. This template does not draft items or run the pilot itself.
What happens after the scale is validated?
The retained pool moves to fielding the full survey, where survey results reporting handles the confidence intervals and subgroup caveats a wider sample supports. A reworded item's psychometrics stay provisional and pending its own revalidation in a future wave rather than carried forward as settled.
Find out which items actually measure what they claim to
Send the drafted item pool and the pilot's response data. The thresholds get fixed first, then every item's fate comes back logged to the exact statistic that decided it.
Edit with AI