Design Critique Framework Template
Four documents and three sheets that hold every item to the requirement it tests, and count the critique that keeps coming back.
Free download · No account needed
Search for a design critique framework and you get a list of question prompts. What problem does this solve, who is it for, what happens when the data is empty, is it accessible. Reasonable questions, and not a framework, because a prompt gives a reviewer nothing to test their answer against. Without that, an hour of six people's time produces a list of preferences, and the designer spends the next day working out which of them came from somebody senior enough to matter.
This pack runs on one rule instead. Every item names the requirement it tests, or it is logged as preference with the reviewer's name against it and it does not block. Preference is not suppressed and it is often right: across the worked quarter in here, 22 of 51 preference items were built anyway. The value is in the label. An item marked preference can be declined in a sentence, where an item claiming a requirement failure has to be investigated first.
The rule produces its own evidence. Merrow, the fictional team in the pack, ran nine reviews and logged 214 items: 138 citing a requirement, 51 preference, 25 questions. In the five reviews where the requirement set went out attached to the invite, 96 of 118 items cited a requirement, at 81.36 per cent. In the four where it did not, 42 of 96 did, at 43.75 per cent. The only variable was whether the thing being tested against was in the room.
What is in the pack
How Feedback Is Bound
The doc to read first, because it is the one move everything else rests on. A facilitation format decides who speaks when. It does not decide what an answer is measured against, so the same room running the same good format produces preference on Tuesday and requirement failures on Thursday depending on whether anybody wrote the requirements down.
Critique Format
The hour itself, short on purpose. Ten minutes of context read aloud, ten minutes of silent review before anyone speaks, thirty minutes of items one at a time with the requirement named, ten minutes of decisions typed on screen. The silent block is the cheapest change available, because the first thing said out loud in a room anchors everything after it.
Feedback Guidance
What reviewers read before their first review. A usable item has three parts: what the design does, which requirement that tests, and why it fails. Accessibility items are the easiest to phrase correctly, because somebody else already wrote the number down: WCAG 2.2 asks for a contrast ratio of at least 4.5:1 on body text. An item saying the contrast looks low is preference. An item saying the secondary label measures 3.1 to 1 against a stated 4.5 to 1 minimum is a requirement item nobody argues with in the room.
Decision Record
Eight fields, written during the review while the room can still disagree with what is being typed. Including the two everyone leaves blank: what was explicitly not decided, and the revisit trigger, which is a condition rather than a date. Objections are recorded, the designer's included, because a decision made over a dated objection teaches the team something.
Review Register and Revision Tracking
One row per review with items logged, items citing a requirement, preference items, questions, blocking items and the named decision owner, plus whether the requirements were attached. Then each item joined to the revision that answered it, with four statuses and no fifth: open, resolved, moved and declined.
Feedback Theme Log
The sheet that changes behaviour, and it counts reviews rather than items. A theme with twenty items in one review is a bad design. A theme with eleven items across four reviews is a system gap. Where the fix belongs takes exactly four values: the design under review, the design system, the requirements template, or product. Two of Merrow's five systemic themes were published accessibility criteria, 25 of the 84 systemic items, and both belonged in a shared component rather than in any of the five designs they were raised against. Target size is the plainest case, since WCAG 2.2 asks for at least 24 by 24 CSS pixels under its minimum criterion. Those route to the accessibility audit remediation pack, the fixed component gets recorded in the design system documentation pack, and inconsistent error wording goes to the content and UX writing pack.
How it works
- 1
Send the designs and the requirements
The artifact under review and whatever the requirements exist as: a spec, acceptance criteria, research findings, a comment thread. Plus your notes from reviews you have already run, because those get classified on the first pass and the counts are usually more persuasive than any argument about process.
- 2
Find out whether the requirements are reviewable
A requirement a reviewer can test says something checkable: entered data survives navigation, a slow query is distinguishable from a failed one within two seconds. A requirement that says the experience should be intuitive cannot be tested, and it will produce preference however well the review is run.
- 3
Run the review and log against requirements
Requirements attached to the invite rather than linked from it, the question stated, settled decisions pulled from the record, one named decision owner and one scribe who is not the designer. Every item classified as requirement, preference or question, with blocking held to three conditions.
- 4
Build the theme log across reviews
Once three or more reviews are logged, group items by the gap rather than by the screen and count how many reviews each theme appeared in. Anything above a third of them is systemic, and it gets priced both ways: rework hours to date against the one-time cost of fixing it upstream.
Frequently asked questions
Is this template free?
Yes. Download the four documents and three sheets as Word and CSV files with nothing to sign up for. Opening the pack as a working space, where the theme log recomputes as reviews are logged and the register totals stay consistent, needs an account.
What format are the downloaded files?
Four Word documents and three CSV sheets in one zip. The documents open in Word, Pages or Google Docs. The sheets open in Excel, Numbers or Google Sheets, and the register and theme log arrive with the worked quarter in them so the column relationships are visible.
Is preference feedback really allowed?
Yes, and labelling it is the point rather than suppressing it. On the worked quarter, 22 of 51 preference items were built anyway. What the label buys is that a designer can decline preference in one sentence, where an item claiming a requirement failure has to be investigated first.
What if we have no written requirements?
Then the setup prompt says so and offers to help write them instead of scheduling a review. Requirements drafted from research notes and support tickets are a fine starting point, and the user flow and journey pack is where the states that need requirements usually get enumerated.
How do you decide a theme is systemic rather than local?
By counting reviews, not items. A theme appearing in a third or more of the reviews is treated as systemic, the threshold is stated on the sheet, and a theme close to it is flagged rather than rounded. Six designs are not independently wrong about loading states.
Does this rank our reviewers or our designers?
No, and it refuses to. Both are computable from what is in the space and both destroy the log, because the next review's items get shaped by the ranking instead of by the requirements. Themes are never attributed to a designer, and items are never counted per reviewer.
How is this different from a critique question checklist?
A checklist gives reviewers prompts. This gives their answers something to be tested against, then keeps the record across reviews so a recurring item stops being relitigated locally. The design handoff checklist is the downstream artifact, once the states have been decided.
Find out which critique your team keeps giving
Send the designs, the requirements and your notes from the reviews you have already run. The first pass classifies those items and tells you whether your requirement set is thick enough to review against.
Run a real critique