River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Concept Testing Plan Template

A concept test with no kill criteria written in advance cannot come back negative, which is why almost none of them do.

Free download  ·  No account needed

This pack writes the criteria before the first session and freezes them. Four to six rows, each carrying what is observed, a confirm number, a kill number, and what literally happens at the kill number. Every response is then scored on what the participant did in the room rather than on whether they said yes. NN/g has been saying since 2001 that self-reported claims and speculation about future behaviour are both unreliable, and a concept test is made almost entirely of the two.

The argument for declaring first does not come from product research. Across 55 large NHLBI trials, 17 of 30 published before 2000 reported a significant benefit against 2 of 25 published after, once the primary outcome had to be registered in advance. The authors read it plainly: almost half of the later trials could have reported something positive had they not declared first. Nothing about the interventions changed.

In the worked test, eleven of twelve said they would use it. Against the criteria frozen fourteen days earlier, two kill conditions fired, one criterion was met at exactly its threshold, and the sample check failed. Built for the fortnight before a build gets committed, not for a research programme. What people already do with a shipped product is a different question, answered by feature adoption review, and a standing weekly habit of talking to customers belongs in the continuous discovery pack.

Twelve sessions, scored against criteria written first

The frozen Decision Criteria, the Response Register behind them, and the Unprompted Problem Log from before the prototype.

Decision Criteria

Illustrative, for a fictional equipment-rental dispatch product called Brambling, testing a self-serve customer booking portal. Frozen fourteen days before the first session, with the consequence agreed at the same time.

CriterionWhat is observedConfirmKillResultVerdict
C-1 The problem is felt without being named for themNames phone booking in their own top three, before the prototype8 of 12under 54 of 12Killed
C-2 Somebody will spend something realNames a pilot date, hands over a customer list, offers an introduction5 of 12under 21 of 12Killed
C-3 The price is not the obstacleTakes $220 a month with no discount ask4 of 120 of 124 of 12Met, at the line
C-4 There is something to stop doingNames the current practice the concept would replace6 of 12under 34 of 12Inconclusive
C-5 The sample is capable of disagreeingRaises the objection we already believe is fatal, unprompted3 of 12under 31 of 12Sample reads agreeable
Stated intent, recorded and not scoredAnswers yes to whether they would use itnot a criterionnot a criterion11 of 12Not scored

Two kill conditions fired independently, one criterion was met at exactly its threshold, and one is inconclusive because it landed between the two numbers rather than being read whichever way suited. C-5 is inverted on purpose and it is the row most likely to be deleted by somebody tidying the sheet: when only one participant will voice the objection everybody already knows about, the finding is that the room could not disagree, which weakens the other four. Against all of that, eleven of twelve said they would use it, recorded on its own row and never counted toward a threshold.

Response Register

One row per participant, scored on the strongest thing they actually did in the session. Recruitment source is a scored column rather than an administrative one.

ParticipantSourceSaid yesC-1C-2C-3Strongest signal
P5, 6 depotsChurned 14 months agoYesYesYesYesCommitted, named a date and two accounts
P2, 2 depotsChose a competitorYesYesNoYesTook the price
P7, 5 depotsChose a competitorYesYesNoYesTook the price
P9, 3 depotsChose a competitorYesNoNoYesTook the price
P11, 2 depotsSucceeding with status quoYesYesNoNoNamed what stops
P1, P3, P8Advisory boardYes0 of 30 of 30 of 3Said yes only
P4, P6, P10No prior relationshipYes0 of 30 of 30 of 3Said yes only
P12, 4 depotsSucceeding with status quoNoNoNoNoSaid no, had used a rival's portal

Eleven of twelve said yes: one committed, three took the price, one named what they would stop, and six gave nothing but the yes. Every group said yes at roughly the same rate and only one group produced anything underneath it. All three advocates said yes and none of them named the problem unprompted, which is what the original recruitment plan of eight advisory-board members would have produced twelve times over. The single commitment came from a company that churned fourteen months ago, and the only decline came from the one participant who had actually used a competitor's portal.

Unprompted Problem Log

The first fifteen minutes of every session, before anything is shown. Open questions about their week, nothing about booking, nothing about the concept.

Problem, in their wordsMentionsIn their top threeBuilt their own workaroundWhat they built
Kit not coming back when it was supposed to9 of 1295Whiteboard, two spreadsheets, a calendar reminder, a WhatsApp group
Arguing about damage that was already there6 of 1242Photos on a personal phone, one printed condition sheet
Deciding which driver goes where tomorrow5 of 1233A magnetic board, a shared calendar, one paid routing tool used for one step
The desk losing its day to the phone4 of 1241A paper day-book
Chasing paperwork before an invoice goes out3 of 1211A folder of scanned delivery notes

The problem the concept was built for came fourth of five, and every one of its four mentions came from somebody who had already left for a competitor or was running the desk personally. The row above it was raised by nine of twelve, which would have passed C-1 outright, and five of those nine had already spent their own time building something for it. A workaround is a specification somebody wrote for themselves under real constraints, so a spreadsheet with conditional formatting and a WhatsApp group with three drivers are two different products. This section is the reason a killed concept test is not a wasted one.

What is in the pack

01

Decision Criteria, frozen

Four to six rows, each with what is observed, a confirm number, a kill number and the consequence that happens at the kill number, agreed before recruitment starts.

02

The criterion that fails when everybody agrees

One inverted row requiring a quarter of participants to raise the objection your team already believes is fatal, because a room that cannot disagree cannot confirm anything either.

03

A screener with a source quota

A cap on advocates, a floor on people who evaluated you and chose somebody else, and at least one person succeeding perfectly well with the status quo.

04

Response Register scored on behaviour

Committed, spent something, named what they would stop, said yes only, said no. Stated intent is recorded on its own row and never counted toward a threshold.

05

Unprompted Problem Log

The fifteen minutes before the prototype appears, ranked by mentions and by how many people had already built their own workaround for each problem.

06

A findings report in a fixed order

Criteria first, stated intent second and labelled unscored, then the consequence in three sentences, because every extra sentence is a place to reintroduce the concept.

How it works

  1. 1

    Answer one question

    What result would make you drop this. Expect the first answer to be a worry and the third to be a number.

  2. 2

    Freeze the criteria

    Confirm threshold, kill threshold and consequence on every row, timestamped and stored before the screener goes out.

  3. 3

    Ask about their week first

    Fifteen minutes of open questions before anything is shown, because after the prototype everybody can name the problem it solves.

  4. 4

    Report against the numbers

    The criteria table, then the yes count labelled unscored, then whatever consequence was agreed, applied rather than reconsidered.

Frequently asked questions

What do I need before this is useful?

A concept, a prototype or sketch in any state, and the number that would make you drop it. That last one is the actual input, and the first answer is usually a worry rather than a threshold. Rough prototypes work better here, because a polished one reads as a decision already made.

Is the point to kill my idea?

No. The point is that a test which cannot come back negative also cannot come back positive, so a confirm result is only worth having if the kill result was reachable. Plenty of concepts survive this. What does not survive is a threshold chosen after the sessions.

Why is one criterion designed to fail when everyone agrees?

Because a sample that will not voice the obvious objection cannot confirm anything either. The row requires at least a quarter of participants to raise, unprompted, whatever your team already believes is fatal. One of twelve raised it in the worked test, which says the room was agreeable rather than the concept safe.

Twelve sessions is not statistically significant.

Correct, and the report never converts a count into a percentage. Twelve distinguishes four from eight from nobody, which is all a count-based criterion needs. Four of twelve named the problem unprompted, never 33 percent of operators, because the moment a count becomes a rate it gets quoted as one.

Who should I recruit?

Not the advisory board, which is where the original plan in the worked test drew eight of twelve. The screener caps advocates at three and requires at least four who evaluated you and chose somebody else. Every advocate said yes and none of them named the problem unprompted.

What happens if the concept is killed?

You get the problem the sessions raised instead. The first fifteen minutes happen before anything is shown, so twelve people rank their own problems unprompted. In the worked test the tested concept came fourth of five; the first was raised by nine of twelve, five of whom had already built a workaround. Those transcripts still code up in research synthesis.

Is this the same as usability testing?

No, and mixing them is how a concept test ends without a decision. Usability testing asks whether people can use a thing already decided on, which a usability test report covers. This asks whether the thing should exist, so the prototype shows the least it can while still being able to fail.

Find out whether your test could come back negative

Send the concept and the number that would make you drop it. What comes back first is the criteria sheet, frozen before anybody is recruited.

Write my kill criteria