River
Y CombinatorBacked by Y Combinator

People & Exec SupportFree

Performance Review Writing With Evidence

A manager's finished review of a strong performer ran 412 words. Exactly one sentence in it described something an outsider could go and verify.

Start here

River reads your own record first and builds an evidence ledger: every line from your 1:1 notes, the goal record, prior reviews and the project tracker, each one marked as something you observed or something you concluded. The draft is then written only from the observation rows. Where a competency has none, River prints the gap instead of writing around it. What the calibration session later reads is exactly this: whether a rating names anything checkable.

A phrase bank makes this worse rather than better. The failure mode of a performance review is not clumsy writing, it is fluent writing with nothing underneath it, and a generator produces the most confident possible version of an unsupported claim. In the worked example the manager's finished draft ran 412 words and contained one checkable sentence, while 22 of the 23 observable outcomes sitting in her own notes never appeared in it at all. The same evidence ledger, aggregated across a team, feeds the review specificity signal in a manager effectiveness comparison.

The second mechanism is an asymmetry. An unsupported compliment is a weak review. An unsupported criticism, appearing in writing for the first time, is the sentence that gets read back to you in a dispute. So River checks every adverse claim for a dated record from before the period closed, keeps the ones that have it, and blocks the ones that do not until you have had the conversation. The evidence discipline is the same one a hiring debrief needs.

Most of a performance review is a conclusion wearing a fact's clothes

Take any review you have written and sort its sentences into two piles. One holds things somebody could go and check: a migration that went live on the committed date, a training week she ran alone after the trainer resigned, two validation checkpoints missed in November. The other holds your conclusions about the person: consistently exceptional, a real asset, goes above and beyond. Both piles read like praise. Only the first survives being asked about, and in the worked example the second pile outnumbered the first by 35 to 23.

The second pile is not dishonest. It is the compression that happens when somebody writes a review in an hour on a Thursday, six months after most of the evidence happened. The information was there and twenty-one weekly notes recorded it. This matters beyond document quality. The EEOC's guidance on appraisal systems notes that evaluations frequently serve as the basis for pay, promotions and terminations. It recommends making sure appraisals rest on employees' actual job performance (EEOC Compliance Manual Section 15). Actual job performance is the first pile.

The asymmetry is what most writing guidance misses. An unsupported compliment costs a weak paragraph. An unsupported criticism is different in kind, especially when the review is the first time the person has heard it, and records behind promotion, demotion and termination decisions have to be retained (29 CFR 1602.14). If nothing else exists, the review is the whole file. So River flags vague passages rather than repairing them, because a fluent version of an unsupported claim is one you cannot defend in the room.

How it works

  1. Send the record

    1:1 notes in whatever state they are in, the goal record, prior reviews, any ticket history.

  2. Lines get sorted

    Each line attributed to a competency and marked observation or conclusion, with its source and date.

  3. The draft gets written

    Only from the observation rows. Gaps are printed rather than filled and vague passages flagged in place.

  4. You resolve the flags

    Name the thing, drop the claim, or record that a competency was not observed this period.

What you get

  • An evidence ledger by competency, every line marked observation or conclusion with its source and date
  • A draft written only from the observation rows, so no sentence exists without something behind it
  • Vague passages flagged inside the draft rather than rewritten, because a fluent unsupported claim is worse
  • Every adverse claim checked for a dated record from before the period closed, and blocked without one
  • The list of work in your own record your draft left out, usually longer than the draft
  • Competencies with no observation recorded as not observed this period, rather than filled in plausibly
  • A count of named outcomes per competency, which is the number calibration is going to read

Common questions

Why not just give me example phrases I can adapt?

Because the phrases are the problem. A review fails when it is fluent and evidences nothing, and a phrase bank is a machine for producing exactly that. The sentence you need is in your 1:1 notes from October, and no amount of borrowed wording will put it there.

Why flag a vague sentence instead of improving it?

Because an improved version of an unsupported claim is more dangerous than the original. It reads specific, so you ship it, and then she asks what you meant and there is no answer. A visible flag forces one of two useful outcomes: name the thing, or drop the claim.

You blocked a criticism I know is true. Why?

Because no record exists that you ever told her, and the word ongoing implies you did. A criticism appearing in writing for the first time in a review is unfair to the person and thin as a record. Have the conversation, write it down, and it is available next cycle.

Is a rating of four out of five evidence?

No. A number is the conclusion you reached, and the ledger asks what you reached it from. This is the same distinction a calibration session applies to the whole cycle: an away-from-middle rating that names nothing checkable gets sent back for substantiation.

What happens when a competency has nothing behind it?

It gets recorded as not observed this period, and the competency is left unrated with the reason written down. That is an honest and useful finding. In the worked example developing others had three conclusions in the record and zero observations, so no paragraph was written for it.

Does this replace what a manager is supposed to think?

No. It will not decide a rating, and it will not tell you whether somebody is good. It finds what your own record already says, separates it from what you concluded, and refuses to write the sentences you cannot stand behind. The judgement stays yours.

Where should the evidence have come from in the first place?

The competencies being rated should be the ones the role was hired against, which the job description defines, and the first record of them starts at onboarding. A review is easier to write when the thing being rated was named before the person started.

Performance Review Writing With Evidence

Fill in the form and your workspace opens with the work already underway.