River
Y CombinatorBacked by Y Combinator
FREE TEMPLATE

Behavioral Interview Question Bank

Four documents and three sheets, including one that scores every question on whether its ratings ever varied and whether they predicted anything.

Free download  ·  No account needed

Discriminating Power by Question, one row

Two numbers decide whether a question keeps its slot

Both computed from your own records. Neither is anybody's opinion of the question.

Modal Share

Of every time this question was asked and rated, the share of ratings that landed on the single most common value. At 90 percent, the question is being asked and everybody is getting the same score, so the slot it consumed produced nothing about the candidate. This needs no performance data at all, only submitted scorecards, so it is available in the first session.

Separation

Among hires who carry a review rating on this same competency, the mean rating of the ones this question scored 3 or 4, minus the mean of the ones it scored 1 or 2. Positive means the question was reading something real. Zero means noise. Negative means it read something and the something was wrong.

Hires With A Criterion Rating

The group size behind the separation figure, printed next to it every time. Under 8 the cell says insufficient outcomes rather than showing a number, because a separation computed on three hires gets quoted and its caveat does not.

Action

Must-ask, keep, retire, or examine the anchor first. A question that inverts with a vague anchor is usually an anchor problem, and rewriting the anchor is cheaper than losing the question.

A question bank only ever grows. Every interviewer adds the question they like, and nobody removes one, because removing a question needs a reason and nobody has a number. So the bank fills with questions that produce the same rating every time. "Tell me about a difficult customer" gets asked, everybody gets a 3, and the slot it consumed produced no information about anybody. The question feels productive, which is exactly why it survives.

This pack computes two numbers per question and lets them decide. Spread is the share of ratings that landed on the most common value, and it runs on submitted scorecards alone. Separation is the mean six-month review rating of the hires a question scored 3 or 4, minus the mean of the ones it scored 1 or 2, on the competency it tests. Frequency of use is deliberately not one of them, because the federal selection guidelines rule out data bearing on the frequency of a procedure's usage as evidence of validity.

In the worked example, Northbeck Instruments had 34 questions across six competencies with 412 rated asks recorded. Nine of the 34 came back at 80 percent modal share or higher, and those nine consumed 118 of the 412 asks: 28.6 percent of every rated question the company asked. Three questions cleared the eight-hire threshold with positive separation, the best at +1.18. The most-loved question in the bank separated at -0.33. Written for recruiters and hiring managers who already run a structured loop and want the questions inside it to earn their slots.

Nine questions, 118 asks, and nothing learned from any of them

The Question Register, the Discriminating Power sheet computed from the Ask Log, and the Ask Log itself.

Discriminating Power by Question

Illustrative for a fictional field service company, Northbeck Instruments, hiring service engineers. Ten of the bank's 34 questions shown, ordered by spread.

QCompAsks4/3/2/1Modal
share
HiresMean if
3-4
Mean if
1-2
SepAction
Q19C5267/9/6/434.6%113.432.25+1.18Must-ask
Q34C3163/6/5/237.5%33.002.00under 8Keep
Q26C2183/7/6/238.9%53.332.50under 8Must-ask on spread
Q11C1194/8/5/242.1%43.332.00under 8Keep as C1 second read
Q14C3275/12/7/344.4%103.172.50+0.67Must-ask
Q03C1316/14/8/345.2%113.252.33+0.92Must-ask
Q23C6299/17/3/058.6%112.673.00-0.33Pull. Anchor first
Q31C4212/18/1/085.7%63.003.00under 8Retire on spread
Q07C2242/21/1/087.5%113.003.000.00Retired 14 Jan
Q18C4221/20/1/090.9%63.173.00under 8Retired 14 Jan
Bank total: 34 questions, 6 competencies, 412 rated asks. Nine questions at 80% modal share or higher, taking 118 asks, which is 28.6% of them. Three of those nine are in the rows above, accounting for 67 asks.

Q18 is the whole finding in one row: 20 of its 22 asks came back a 3. Q23 is the harder one. Its spread is fine and its direction is wrong, and the anchor rewards fluency about the future rather than anything that happened, so the anchor gets examined before the question does.

Question Register

Both anchor ends written is a condition of entry, not a nice-to-have. Retirement is a dated row change with the replacement named.

QCompetencyWtQuestion as askedAnchors
4 / 1
StatusReplaced
by
Q19C525Walk me through the last calibration drift you caught before the customer noticed. How did you catch it and what did you do?yes / yesmust-askn/a
Q03C125Tell me about the last time a customer described a symptom and the actual cause was somewhere else entirely. What did you check, in what order?yes / yesmust-askn/a
Q14C320You have four open tickets and two are inside their SLA window. Tell me about the last time you were actually in that position and what you did.yes / yesmust-askn/a
Q26C215A site manager who signs the invoice asks why the visit took two days instead of one. Tell me about a real time you had that conversation.yes / yesmust-askn/a
Q07C215How would you explain what you do to somebody with no technical background?yes / yesretiredQ26
Q18C410Tell me about a difficult customer.no / noretiredQ41
Q31C410How do you stay calm under pressure?no / noretiredQ41
Q41C410Tell me about a time a site contact escalated over your head. What had happened, what did you do next, and how did it end?yes / yescandidate, 0 asksn/a
Q23C65Where do you want to be in five years?withdrawnpulled from must-askn/a

C4 lost both of its questions in one pass. It is now carried by one drafted question with zero recorded asks, and the C4 rating anchor is marked on hold until Q41 clears 12 asks. That is an uncomfortable finding and the correct one: a live anchor over a competency with no measured read is how a bank starts producing ratings nobody can defend.

Ask Log

One row per question per candidate. Ratings without a written evidence cell are recorded as missing, not as a low score.

DateCandidateQInterviewerRatedEvidence recordedHired6-month C-rating
11 SepC-0412Q19Ferreira4went looking unprompted, weekly logyes4
11 SepC-0412Q18Ferreira3noneyes4
18 SepC-0431Q19Okonkwo2found it on the customer callyes2
18 SepC-0431Q07Okonkwo3clear analogy, no exampleyes3
02 OctC-0455Q23Vasquez4named a path and a timelineyes2
02 OctC-0455Q14Vasquez2no sequencing rule offeredyes2
09 OctC-0478Q18Ferreira3nonenon/a
09 OctC-0478Q31Bhatt3nonenon/a

The two Q18 rows are the pattern in miniature: same question, same rating, no evidence either time, and one candidate who went on to a 4 while the other was rejected. Both rows say 3. The log is also where the criterion comes from, and it is the per-competency review rating rather than an overall score, because the guidelines want criteria that represent the [important job duties developed from the review of job information](https://www.ecfr.gov/current/title-29/subtitle-B/chapter-XIV/part-1607/section-1607.14).

What is in the pack

01

Discriminating Power by Question

Modal share, the rating distribution behind it, the separation figure, and the number of hires that figure rests on. Under eight hires the cell says insufficient outcomes rather than showing a number, because the number gets quoted and the caveat does not.

02

Question Register

One row per question with its competency, its weight, the question as it is actually asked, whether both anchor ends are written, and its status. Retirements stay on the sheet as dated rows with the replacement named.

03

Ask Log

One row per question per candidate: what was asked, what it was rated, what the interviewer wrote as evidence, and the post-hire rating where one exists. Every other number in the pack recomputes from this sheet.

04

Question Bank by Competency

The surviving questions grouped by competency, each with its follow-up probes and the specific thing that separates a 4 from a 3. Competencies with no qualifying question are shown as open gaps rather than quietly omitted.

05

Rating Anchors

A four-point scale with both ends written for every competency that has a qualifying question. Where a competency lost its last one, the anchor is marked on hold instead of left live over a gap.

06

Interviewer Guidance

How to run the probes, why evidence gets written before the rating, and what not to ask. It also says not to read other panelists' scorecards first, which is the cheapest correction available to any loop.

07

How a Question Earns a Slot

The four conditions and the two thresholds, with the reasoning for each. Twelve recorded asks before a spread verdict, eight rated hires before a separation figure, and popularity excluded on purpose.

How it works

  1. 1

    Send the ask record

    Submitted scorecards with the question text on them, an ATS scorecard export, or interviewer guides with ratings written on them. What somebody rated is the record of what they asked, so scorecards beat a list of questions every time.

  2. 2

    Spread is scored first

    River normalises the question text so the same question asked in four wordings collapses into one row, then reports modal share per question. This half needs no performance data, so it lands in the first session.

  3. 3

    Separation, where outcomes support it

    Send per-competency review ratings for people hired into this role family and River joins them to the log. Competencies under eight rated hires come back as insufficient rather than as a weak number.

  4. 4

    Retire and replace in one pass

    Every retirement carries a replacement drafted against the specific failure that retired the original, with both anchor ends written before it enters the bank as a candidate.

Frequently asked questions

Why not just download a list of behavioral interview questions?

Because a list is what every bank started as, and a list is what degrades. The lists on page one are ordered by competency and nothing else, so a question that produced the same rating on 22 consecutive asks renders identically to one that separates hires by a full rating point. The two numbers here are what tell them apart.

We have no post-hire performance data. Is this useless?

No, and it is worth being specific about why. Spread runs on submitted scorecards alone, so the entire retirement case for the nine flat questions in the worked example needed no outcome data at all. Separation is the second half and it arrives later. Start with spread and load review ratings when a full cycle has closed.

Why eight hires before you compute separation?

Because a difference of means on three hires is one person's bad quarter. Eight is a working floor rather than a statistical claim, and the pack prints the group size beside every figure so a reader can judge it themselves. The guidelines leave sample size to the user, which is a polite way of saying small samples do not become evidence by being written down.

One of our best questions came back with negative separation. Do we bin it?

Examine the anchor first. In the worked example, "where do you want to be in five years" separated at -0.33 across 11 hires, and its anchor rewarded fluency about the future rather than anything that had happened. That is an anchor defect wearing a question's clothes. It was pulled from must-ask with the anchor flagged, not retired.

Why does a retired question stay on the sheet?

Because it contributed to hiring decisions, and personnel records having to do with hiring must be preserved for at least a year from the date of the record or the personnel action, longer once a charge is filed. Deleting the row deletes the evidence that the bank was maintained on evidence. Retirement is a dated status change with the reason written out.

How does this fit with the rest of the hiring process?

It sits underneath the loop. The interview loop pack decides which competency each interviewer owns and how many minutes it gets, and the competencies themselves come off the job description and requirement register. This space maintains the questions inside those slots. After a loop runs, debrief synthesis reads the scorecards it produced, and those scorecards are the material this space loads next time.

What is in the download, and what does Edit with AI add?

Four Word documents and three CSV sheets in one zip, free and with no account needed. Edit with AI installs the same pack as a private Space and scores your own bank: your scorecards, your competencies, your outcomes. If the constraint is application volume rather than question quality, resume screening sits upstream of all of this.

Find out which of your questions is doing nothing

Send your submitted scorecards and your competency list. The first thing back is the modal share on every question and what share of your interview time went to the ones that never varied.

Edit with AI