Behavioral Interview Question Bank
Four documents and three sheets, including one that scores every question on whether its ratings ever varied and whether they predicted anything.
Free download · No account needed
Discriminating Power by Question, one row
Two numbers decide whether a question keeps its slot
Both computed from your own records. Neither is anybody's opinion of the question.
Modal Share
Of every time this question was asked and rated, the share of ratings that landed on the single most common value. At 90 percent, the question is being asked and everybody is getting the same score, so the slot it consumed produced nothing about the candidate. This needs no performance data at all, only submitted scorecards, so it is available in the first session.
Separation
Among hires who carry a review rating on this same competency, the mean rating of the ones this question scored 3 or 4, minus the mean of the ones it scored 1 or 2. Positive means the question was reading something real. Zero means noise. Negative means it read something and the something was wrong.
Hires With A Criterion Rating
The group size behind the separation figure, printed next to it every time. Under 8 the cell says insufficient outcomes rather than showing a number, because a separation computed on three hires gets quoted and its caveat does not.
Action
Must-ask, keep, retire, or examine the anchor first. A question that inverts with a vague anchor is usually an anchor problem, and rewriting the anchor is cheaper than losing the question.
A question bank only ever grows. Every interviewer adds the question they like, and nobody removes one, because removing a question needs a reason and nobody has a number. So the bank fills with questions that produce the same rating every time. "Tell me about a difficult customer" gets asked, everybody gets a 3, and the slot it consumed produced no information about anybody. The question feels productive, which is exactly why it survives.
This pack computes two numbers per question and lets them decide. Spread is the share of ratings that landed on the most common value, and it runs on submitted scorecards alone. Separation is the mean six-month review rating of the hires a question scored 3 or 4, minus the mean of the ones it scored 1 or 2, on the competency it tests. Frequency of use is deliberately not one of them, because the federal selection guidelines rule out data bearing on the frequency of a procedure's usage as evidence of validity.
In the worked example, Northbeck Instruments had 34 questions across six competencies with 412 rated asks recorded. Nine of the 34 came back at 80 percent modal share or higher, and those nine consumed 118 of the 412 asks: 28.6 percent of every rated question the company asked. Three questions cleared the eight-hire threshold with positive separation, the best at +1.18. The most-loved question in the bank separated at -0.33. Written for recruiters and hiring managers who already run a structured loop and want the questions inside it to earn their slots.
What is in the pack
Discriminating Power by Question
Modal share, the rating distribution behind it, the separation figure, and the number of hires that figure rests on. Under eight hires the cell says insufficient outcomes rather than showing a number, because the number gets quoted and the caveat does not.
Question Register
One row per question with its competency, its weight, the question as it is actually asked, whether both anchor ends are written, and its status. Retirements stay on the sheet as dated rows with the replacement named.
Ask Log
One row per question per candidate: what was asked, what it was rated, what the interviewer wrote as evidence, and the post-hire rating where one exists. Every other number in the pack recomputes from this sheet.
Question Bank by Competency
The surviving questions grouped by competency, each with its follow-up probes and the specific thing that separates a 4 from a 3. Competencies with no qualifying question are shown as open gaps rather than quietly omitted.
Rating Anchors
A four-point scale with both ends written for every competency that has a qualifying question. Where a competency lost its last one, the anchor is marked on hold instead of left live over a gap.
Interviewer Guidance
How to run the probes, why evidence gets written before the rating, and what not to ask. It also says not to read other panelists' scorecards first, which is the cheapest correction available to any loop.
How a Question Earns a Slot
The four conditions and the two thresholds, with the reasoning for each. Twelve recorded asks before a spread verdict, eight rated hires before a separation figure, and popularity excluded on purpose.
How it works
- 1
Send the ask record
Submitted scorecards with the question text on them, an ATS scorecard export, or interviewer guides with ratings written on them. What somebody rated is the record of what they asked, so scorecards beat a list of questions every time.
- 2
Spread is scored first
River normalises the question text so the same question asked in four wordings collapses into one row, then reports modal share per question. This half needs no performance data, so it lands in the first session.
- 3
Separation, where outcomes support it
Send per-competency review ratings for people hired into this role family and River joins them to the log. Competencies under eight rated hires come back as insufficient rather than as a weak number.
- 4
Retire and replace in one pass
Every retirement carries a replacement drafted against the specific failure that retired the original, with both anchor ends written before it enters the bank as a candidate.
Frequently asked questions
Why not just download a list of behavioral interview questions?
Because a list is what every bank started as, and a list is what degrades. The lists on page one are ordered by competency and nothing else, so a question that produced the same rating on 22 consecutive asks renders identically to one that separates hires by a full rating point. The two numbers here are what tell them apart.
We have no post-hire performance data. Is this useless?
No, and it is worth being specific about why. Spread runs on submitted scorecards alone, so the entire retirement case for the nine flat questions in the worked example needed no outcome data at all. Separation is the second half and it arrives later. Start with spread and load review ratings when a full cycle has closed.
Why eight hires before you compute separation?
Because a difference of means on three hires is one person's bad quarter. Eight is a working floor rather than a statistical claim, and the pack prints the group size beside every figure so a reader can judge it themselves. The guidelines leave sample size to the user, which is a polite way of saying small samples do not become evidence by being written down.
One of our best questions came back with negative separation. Do we bin it?
Examine the anchor first. In the worked example, "where do you want to be in five years" separated at -0.33 across 11 hires, and its anchor rewarded fluency about the future rather than anything that had happened. That is an anchor defect wearing a question's clothes. It was pulled from must-ask with the anchor flagged, not retired.
Why does a retired question stay on the sheet?
Because it contributed to hiring decisions, and personnel records having to do with hiring must be preserved for at least a year from the date of the record or the personnel action, longer once a charge is filed. Deleting the row deletes the evidence that the bank was maintained on evidence. Retirement is a dated status change with the reason written out.
How does this fit with the rest of the hiring process?
It sits underneath the loop. The interview loop pack decides which competency each interviewer owns and how many minutes it gets, and the competencies themselves come off the job description and requirement register. This space maintains the questions inside those slots. After a loop runs, debrief synthesis reads the scorecards it produced, and those scorecards are the material this space loads next time.
What is in the download, and what does Edit with AI add?
Four Word documents and three CSV sheets in one zip, free and with no account needed. Edit with AI installs the same pack as a private Space and scores your own bank: your scorecards, your competencies, your outcomes. If the constraint is application volume rather than question quality, resume screening sits upstream of all of this.
Find out which of your questions is doing nothing
Send your submitted scorecards and your competency list. The first thing back is the modal share on every question and what share of your interview time went to the ones that never varied.
Edit with AI