Concept Testing Plan Template
A concept test with no kill criteria written in advance cannot come back negative, which is why almost none of them do.
Free download · No account needed
This pack writes the criteria before the first session and freezes them. Four to six rows, each carrying what is observed, a confirm number, a kill number, and what literally happens at the kill number. Every response is then scored on what the participant did in the room rather than on whether they said yes. NN/g has been saying since 2001 that self-reported claims and speculation about future behaviour are both unreliable, and a concept test is made almost entirely of the two.
The argument for declaring first does not come from product research. Across 55 large NHLBI trials, 17 of 30 published before 2000 reported a significant benefit against 2 of 25 published after, once the primary outcome had to be registered in advance. The authors read it plainly: almost half of the later trials could have reported something positive had they not declared first. Nothing about the interventions changed.
In the worked test, eleven of twelve said they would use it. Against the criteria frozen fourteen days earlier, two kill conditions fired, one criterion was met at exactly its threshold, and the sample check failed. Built for the fortnight before a build gets committed, not for a research programme. What people already do with a shipped product is a different question, answered by feature adoption review, and a standing weekly habit of talking to customers belongs in the continuous discovery pack.
What is in the pack
Decision Criteria, frozen
Four to six rows, each with what is observed, a confirm number, a kill number and the consequence that happens at the kill number, agreed before recruitment starts.
The criterion that fails when everybody agrees
One inverted row requiring a quarter of participants to raise the objection your team already believes is fatal, because a room that cannot disagree cannot confirm anything either.
A screener with a source quota
A cap on advocates, a floor on people who evaluated you and chose somebody else, and at least one person succeeding perfectly well with the status quo.
Response Register scored on behaviour
Committed, spent something, named what they would stop, said yes only, said no. Stated intent is recorded on its own row and never counted toward a threshold.
Unprompted Problem Log
The fifteen minutes before the prototype appears, ranked by mentions and by how many people had already built their own workaround for each problem.
A findings report in a fixed order
Criteria first, stated intent second and labelled unscored, then the consequence in three sentences, because every extra sentence is a place to reintroduce the concept.
How it works
- 1
Answer one question
What result would make you drop this. Expect the first answer to be a worry and the third to be a number.
- 2
Freeze the criteria
Confirm threshold, kill threshold and consequence on every row, timestamped and stored before the screener goes out.
- 3
Ask about their week first
Fifteen minutes of open questions before anything is shown, because after the prototype everybody can name the problem it solves.
- 4
Report against the numbers
The criteria table, then the yes count labelled unscored, then whatever consequence was agreed, applied rather than reconsidered.
Frequently asked questions
What do I need before this is useful?
A concept, a prototype or sketch in any state, and the number that would make you drop it. That last one is the actual input, and the first answer is usually a worry rather than a threshold. Rough prototypes work better here, because a polished one reads as a decision already made.
Is the point to kill my idea?
No. The point is that a test which cannot come back negative also cannot come back positive, so a confirm result is only worth having if the kill result was reachable. Plenty of concepts survive this. What does not survive is a threshold chosen after the sessions.
Why is one criterion designed to fail when everyone agrees?
Because a sample that will not voice the obvious objection cannot confirm anything either. The row requires at least a quarter of participants to raise, unprompted, whatever your team already believes is fatal. One of twelve raised it in the worked test, which says the room was agreeable rather than the concept safe.
Twelve sessions is not statistically significant.
Correct, and the report never converts a count into a percentage. Twelve distinguishes four from eight from nobody, which is all a count-based criterion needs. Four of twelve named the problem unprompted, never 33 percent of operators, because the moment a count becomes a rate it gets quoted as one.
Who should I recruit?
Not the advisory board, which is where the original plan in the worked test drew eight of twelve. The screener caps advocates at three and requires at least four who evaluated you and chose somebody else. Every advocate said yes and none of them named the problem unprompted.
What happens if the concept is killed?
You get the problem the sessions raised instead. The first fifteen minutes happen before anything is shown, so twelve people rank their own problems unprompted. In the worked test the tested concept came fourth of five; the first was raised by nine of twelve, five of whom had already built a workaround. Those transcripts still code up in research synthesis.
Is this the same as usability testing?
No, and mixing them is how a concept test ends without a decision. Usability testing asks whether people can use a thing already decided on, which a usability test report covers. This asks whether the thing should exist, so the prototype shows the least it can while still being able to fail.
Find out whether your test could come back negative
Send the concept and the number that would make you drop it. What comes back first is the criteria sheet, frozen before anybody is recruited.
Write my kill criteria