River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Scale Validation Procedure Template

Two documents and four sheets that connect each item's fate, kept, reworded or dropped, to the exact statistic that decided it

Free download  ·  No account needed

Item Revision Log

Change Readiness Scale, Highmark Logistics pilot

ItemStatistic that firedDecision
C4Item-total r=.20; alpha rises to .84 if dropped; loading .34, communality .13Drop
E5Item-total r=.44 and alpha-if-excluded=.82 both pass; loading .54 lands on the wrong factorReword, hold for revalidation
B4Item-total r=.59; alpha falls, not rises, if excludedRetain

C4 fails all three sheets. E5 passes two and fails the third, the case this log exists to catch. B4 is reviewed and stays.

A scale-validation tutorial states the actual decision rule in one row of one table. A corrected item-total correlation under .30, or a rising Cronbach's alpha, is reason enough to flag an item, a rule a 2018 protocol in Health Psychology and Behavioral Medicine states inside one study's results narrative. A 2024 best-practice review adds a second test: whether an item's factor loading lands on a different construct once every item runs against every factor together. Neither names a place to log an item that clears the first test and fails the second.

This space fixes those checks as three standing thresholds before a single pilot number is read, then logs every item's fate to the exact statistic that decided it. Item-total and reliability are both computed against an item's own subscale in isolation, so an item can clear both while still measuring the wrong construct; only the factor-structure check tests it against everything else in the pool at once. It does not draft the item pool itself. A dedicated instrument-design pack does that, tracing every item to a construct before a single response exists.

In the worked example, Highmark Logistics pilots an 18-item Change Readiness Scale on 238 employees ahead of a warehouse system changeover. C4 fails every sheet: an item-total correlation of .20, an alpha that rises to .84 without it, and a factor loading of just .34 on its own construct. E5 clears item-total and reliability cleanly but loads .54 on the wrong factor against .42 on its own, the pattern only the third check catches, so it is reworded and held for revalidation rather than treated as finished.

Every sheet in the pack

Item Statistics, Reliability by Subscale, Factor Structure, and Item Revision Log.

Item Statistics

Mean, SD and corrected item-total correlation for every item, computed against its own subscale. Change Readiness Scale, Highmark Logistics pilot, N=238.

ItemSubscaleMeanSDItem-total rFlag
C1Cognitive3.000.920.62
C2Cognitive3.030.920.64
C3Cognitive3.010.860.63
C4Cognitive2.970.920.20Below .30
C5Cognitive3.000.880.62
C6Cognitive3.030.890.59
E1Emotional2.990.890.67
E5Emotional2.980.900.44
B4Behavioral3.030.920.59

Every item above clears the .30 item-total floor, including E5 and B4. The other nine items across Emotional and Behavioral Readiness run from .55 to .69; the full sheet holds all eighteen.

Reliability by Subscale

Cronbach’s alpha as drafted, the item whose removal would raise it most, and the revised alpha once that item is actually dropped or kept.

SubscalekAlpha (drafted)Highest alpha-if-excludedAlpha if excludedRevised
Cognitive Readiness60.79C40.840.84 (5 items, C4 dropped)
Emotional Readiness60.81E50.820.81 (6 items, E5 reworded)
Behavioral Readiness60.86B40.850.86 (6 items, all retained)

C4 is the only item in the pool whose removal raises alpha above the as-drafted value. E5’s gain is marginal and would not flag it alone; Factor Structure is what actually catches E5.

Factor Structure

Every item tested against every factor at once, not just its own. A loading above .30 on a non-home factor is the one worth reading past the header row.

ItemHomeCognitiveBehavioralEmotionalCommunalityFlag
C1Cognitive0.740.080.070.55
C2Cognitive0.79-0.010.010.62
C3Cognitive0.78-0.010.010.61
C4Cognitive0.34-0.02-0.080.13Weak on own factor too
C5Cognitive0.79-0.04-0.030.63
C6Cognitive0.720.050.000.53
E1Emotional0.05-0.020.790.62
E5Emotional-0.080.540.420.48Loads on Behavioral, not home factor
B4Behavioral0.040.700.000.49

E5 is the item that clears Item Statistics and Reliability and fails here: .54 on Behavioral against .42 on its own Emotional factor, a .12 gap past the .10 minor-difference line. The other nine items land on their own home factor; the full sheet holds all eighteen.

Item Revision Log

Which threshold fired, the exact number, and the decision, for every flagged item.

ItemStatistic that firedDecision
C4Item-total .20; alpha rises to .84 if dropped, from .79; loading .34, communality .13Drop
E5Item-total .44 and alpha-if-excluded .82 both pass; loading .54 on Behavioral versus .42 on its own factorReword, hold for revalidation
B4Item-total .59 clears .30; alpha falls to .85 if excluded, from .86, the smallest margin in its subscaleRetain
Remaining 15 itemsItem-total .55 or above; no alpha-if-excluded above its subscale’s drafted alpha; no loading outside its home factor by more than a minor differenceRetain as drafted

Three items get their own row because a threshold fired. Fifteen clear every sheet and get one summary row rather than fifteen identical ones.

What's in the pack

01

Method Note

The three decision thresholds, item-total, reliability and factor structure, fixed and named before a single item statistic from the actual pilot is computed.

02

Item Statistics sheet

Every item's mean, standard deviation and corrected item-total correlation against its own subscale, with anything under .30 flagged in the same row.

03

Reliability by Subscale sheet

Cronbach's alpha as drafted for each subscale, the item whose removal would raise it most, and the revised alpha once that item is actually dropped or kept.

04

Factor Structure sheet

Every item's loading on every factor at once, not just its own, with communality and a flag for a loading that lands outside its home factor by more than a minor gap.

05

Item Revision Log sheet

Which threshold fired for every flagged item, the exact number that fired it, and the decision, drop, reword or retain, with one summary row for everything that cleared all three.

06

Validation Narrative

The pattern read in prose once every item's fate is settled: what was dropped, what was reworded and why, and what one pilot does not establish on its own.

07

An item's fate traces to every sheet that tested it, not just the one that passed

The standing space rule every prompt reads first, including why two passing numbers do not rule out a factor-structure problem.

How to use it

  1. 1

    Open in River, or download it

    Open the pack in River and the agent fixes the thresholds and runs the checks against your own pilot data, or download the two documents and four CSV sheets instantly, filled in with the worked example.

  2. 2

    Send the item pool and the pilot data

    The drafted items grouped by the construct or subscale each one was written for, and the pilot's response data in rows and columns, however it was collected.

  3. 3

    Thresholds get fixed before a number is read

    Item-total, reliability and factor-structure thresholds go into Method Note before a single statistic from the actual pilot is computed, so no threshold gets picked after seeing which items would fail it.

  4. 4

    Every item's fate gets logged, then read as one pattern

    Item Revision Log states which threshold fired and what the number was for every flagged item, and Validation Narrative reads what the full pattern does and does not establish.

Frequently asked questions

Is this template free?

Yes. Download the whole pack as a Word document and CSV sheets with no signup and no credit card, filled in for the Highmark Logistics worked example. Edit with AI is a separate, optional path that fixes the thresholds and runs the checks against your own pilot data. The template library holds the rest of the packs.

What format are the downloaded files?

Word documents for Method Note and Validation Narrative, and CSV for the four sheets, all in one zip. They open in Word, Pages, Google Docs, Excel, Numbers and Sheets with nothing to convert. Inside River the same content opens as live Docs and Sheets.

Does this draft the survey items, or decide what the scale measures?

No. It assumes the item pool and the pilot data already exist and decides which items keep their place, get reworded, or get dropped. A dedicated instrument-design pack traces every item to a construct and a planned analysis before a single response is collected, which is a separate, earlier step.

How small can the pilot sample be?

The sources behind this template's thresholds treat a participant-to-item ratio between 5 to 1 and 10 to 1 as commonly followed for a factor analysis. The worked example pilots 238 respondents against 18 items, a 13 to 1 ratio. Below that range, item statistics and reliability still run, but the factor-structure check becomes the least trustworthy of the three.

What happens to an item that clears two checks and fails the third?

That is the case this template is built to catch. Item-total correlation and reliability are both computed against an item's own subscale in isolation, so an item can pass both while its factor loading lands on a different construct once tested against the whole pool. The item gets reworded and held for revalidation, not waved through.

Does it run the statistics for me, or do I need the pilot data already collected?

It needs your pilot's response data already in rows and columns, one column per item. If the export still carries a survey platform's raw formatting, the export preparation tool handles the platform quirks first. This template does not draft items or run the pilot itself.

What happens after the scale is validated?

The retained pool moves to fielding the full survey, where survey results reporting handles the confidence intervals and subgroup caveats a wider sample supports. A reworded item's psychometrics stay provisional and pending its own revalidation in a future wave rather than carried forward as settled.

Find out which items actually measure what they claim to

Send the drafted item pool and the pilot's response data. The thresholds get fixed first, then every item's fate comes back logged to the exact statistic that decided it.

Edit with AI