River
Y CombinatorBacked by Y Combinator

People & Exec SupportFree

Hiring Debrief Interview Scorecard Synthesis

A panel of four produced 34 lines of feedback, one competency nobody tested, and a 3.0 average that decides nothing.

Start here

River's debrief synthesis takes the scorecards and written feedback in whatever shape they arrived and attributes every line to the competency it actually evidences. Back comes a sheet with one row per competency and one column per interviewer, a written debrief naming where the panel agrees, where it genuinely conflicts, and the single question each conflict turns on. It never averages the ratings, because the average of a strong yes and a considered no is the least useful number a panel can produce.

Most disagreement in a debrief is not disagreement. Two interviewers reach opposite verdicts because they tested different things and each generalized from what they saw. Re-attributing the evidence separates those cases from the real ones. In the worked example, four interviewers produced three apparent conflicts and only one survived the pass: the hiring manager and the peer consultant had both tested go-live ownership and had contradictory answers about who made the cutover call. The competency list comes from the job description pack rather than being invented here.

Written for recruiters running the debrief and hiring managers who have to defend the outcome later, plus anyone who has sat in a meeting where the loudest interviewer set the decision. Reach for it before the debrief rather than after, because the coverage gap is the finding you cannot fix once everybody has left. If the shortlist itself is the problem, resume screening on evidence is the earlier step, and interview transcript synthesis handles recordings.

Most of what a panel writes down is a conclusion

Read any set of scorecards and sort the lines into two piles. One pile holds things somebody watched: when asked about the cutover date, she said the partner made the call. The other holds verdicts: did not seem strategic, lacks executive presence, strong culture fit. In the worked example the split was 13 observations against 21 conclusions with nothing behind them. The conclusions are what a debrief argues about, and they are the pile that cannot be checked, reconciled, or defended six months later.

The federal guidelines on selection procedures are unusually direct about this. Under the Uniform Guidelines, a content validity strategy is not appropriate for a procedure that purports to measure traits or constructs, and the list of examples given is intelligence, aptitude, personality, commonsense, judgment, leadership. An interview is a selection procedure. It is defensible to the extent it samples the behavior of the job. Every line reporting a construct rather than a behavior sits outside what the procedure can support.

The second finding is a hole rather than a conflict. Rank the competencies by how many interviewers actually tested each one and the distribution is never flat. In the worked example three of the four tested go-live ownership, one tested the data work, and nobody tested training delivery at all, on a role where every implementation ends in a training week. A panel that has covered one competency three times and another zero times has not assessed the candidate, and averaging four verdicts hides exactly that.

How it works

  1. Send the scorecards

    Whatever shape they arrived in: the applicant tracking export, Google Docs, email, and the notes nobody submitted.

  2. Lines get sorted

    Each one attributed to a competency and marked observation or conclusion. Unassignable lines are listed separately.

  3. Read the conflicts

    Which ones dissolve into two people testing different things, which one is real, and the question it turns on.

  4. Close the gaps

    The competencies nobody covered, with the shortest thing that would test each before a decision gets made.

What you get

  • Every feedback line attributed to the competency it evidences, with the interviewer who supplied it
  • Observations separated from conclusions, so a verdict with nothing behind it stops carrying weight
  • Coverage by competency across interviewers, including the ones nobody tested and the ones tested three times
  • Apparent conflicts split into coverage artifacts and genuine contradictions on the same competency
  • One specific question per real conflict, with who can answer it and how long that takes
  • A written debrief that never averages the ratings and says what is still outstanding

Common questions

Our scorecards have numeric ratings. Why not just average them?

Because the average destroys the only information the panel produced. A strong yes and a considered no average to a hire, and so do four lukewarm maybes, and those are opposite situations. The numbers get reported individually next to the evidence each one rests on, and where two interviewers rated the same competency differently the sheet shows why.

What if we never agreed what each interviewer was testing?

Then the coverage finding is the main output. Competencies get derived from the role and every line assigned to the nearest one, which reliably shows three people testing one thing and two competencies untested. That repeats on every candidate, and it surfaces in the funnel as a slow late stage that recruiting funnel analysis can find but not explain.

Does it tell me whether to hire?

It tells you what is still outstanding and what would settle it. In the worked example the answer was two items, both closable inside 48 hours: one reference question about who made the cutover call, and a 15-minute walkthrough to test the competency nobody had covered. A verdict on incomplete evidence is the thing to avoid.

One of our interviewers writes two sentences. Is that a problem?

It is a finding, and it gets reported as one. Two sentences of conclusion contribute nothing to a competency row, so that interviewer's verdict stands on nothing the panel can inspect. Naming it once tends to fix it, because the person can see their own column next to a colleague's on the same sheet.

Is not a culture fit really unusable?

It is unusable as written, and the pass says so rather than deleting it. The useful move is to ask what the interviewer actually watched. Sometimes culture fit resolves into something observable, usually one specific interaction. When it resolves into nothing, that is the answer, and the debrief records it that way.

How is this different from what our applicant tracking system already does?

The system stores the scorecards and computes an average. It has no model of which competency a line of free text is evidence about, so it cannot tell a coverage gap from a conflict, and it will report a 3.0 for a panel that split four ways. Synthesis is the step it was never built to do.

Do I need to keep the debrief afterwards?

Keep it. Hiring records are preserved under EEOC recordkeeping rules for a year from the record or the personnel action, whichever is later, and longer once a charge is filed. A debrief carrying the observation behind each verdict is a better record than four scorecards and an average.

Hiring Debrief Interview Scorecard Synthesis

Fill in the form and your workspace opens with the work already underway.