River
Y CombinatorBacked by Y Combinator

Sales & PartnershipsFree

Sales Activity Analysis Against Win Rates

Activity counts against outcomes, normalized for how long each deal stayed open, so the metric on your dashboard gets tested rather than assumed.

Start here

River returns a sheet, a chart and a short document. The sheet holds every activity metric you named plus the ones worth testing, computed per open day and inside a fixed early window, with won and lost side by side. The chart plots each metric against outcome so the flat ones look flat. The document says which metrics to coach toward, which to score deals with, and which to stop reporting. The metric already on your dashboard is tested first and named either way.

The two corrections are what change the answer. Raw activity counts reward deals that lived longer, and won deals live roughly three times as long as lost ones, so a two-fold activity gap can be entirely duration. Dividing by days open removes that. Measuring only the first three weeks, on deals that were still open at three weeks, removes the other half of the problem: activity that happened because the deal was already going well. Neither correction is complicated and neither is on the dashboard.

Written for the sales manager about to set next quarter's activity targets, the revenue operations analyst asked to prove the current ones work, and the founder who inherited a dashboard nobody built on purpose. Run it before the targets are set rather than after, since a number already written into a comp plan is far harder to retire than one still being drafted. Which stage to aim the finding at is the conversion diagnostic, and the coaching that follows is the call coaching pack.

Why activity counts favour the deals that lasted

Counting events over unequal exposure is a solved problem everywhere except sales dashboards. Workplace injury reporting cannot compare a plant that worked 400,000 hours against one that worked 90,000, so the rule normalizes: total injuries times 200,000, divided by hours worked by all employees. Days open is the hours worked. Ranking won deals against lost ones on raw activity counts is the same error as ranking those two plants on injury counts alone, and it picks the same wrong plant.

The second failure is testing everything and reporting whatever cleared. Regulators are explicit about it: as the number of endpoints or analyses increases, the Type I error rate can increase well beyond 2.5 percent, which is why the Bonferroni correction divides the threshold by the number of tests. Fourteen activity metrics tested at the usual threshold yield 0.7 spurious passes from noise alone. The sheet prints how many were tested and holds every survivor to the divided threshold.

Marloway Systems, twelve months, 640 opportunities and 520 resolved. Won deals average 31.4 touches against 11.2 for lost, a 2.8x gap that reads as proof of something. They also stay open 94 days against 31. Per open day the gap inverts to 0.334 against 0.361. Inside the first 21 days it is 7.9 against 7.4. Their twelve-touch target separates win rates by 1.4 points, and reaching three buyer contacts separates them by 35.6.

How it works

  1. Rebuild the clock

    Days open for every opportunity from its created and closed dates, since every rate below divides by it.

  2. Test what you manage

    Your own targets first, raw and per open day, with the distance between those two answers reported.

  3. Set the landmark

    Recount inside a fixed early window, keeping only deals that were still open at the end of it.

  4. Sort by control

    Survivors split into what a rep can change, what the buyer decides, and what to stop reporting.

What you get

  • Every metric you already manage, tested first and named in the output either way
  • Activity per open day rather than per deal, so duration stops doing the work
  • A fixed early window on deals that survived it, before an outcome could cause the activity
  • The count of metrics tested and a threshold divided by it, printed side by side
  • Metrics split by who controls them, since a buyer behaviour is a score and not a target
  • The sample size each finding would need to be proved, against your actual deal volume

Common questions

Surely more activity helps. Why would it not?

It might, and this does not assume otherwise. What it refuses to accept is the raw comparison, because won deals in the worked example stay open 94 days against 31 for lost ones. Three times the window produces roughly three times the activity on its own, before anybody sells anything.

Why cut the window at three weeks?

Because activity after that point is partly an effect rather than a cause. A deal going well generates meetings, and counting those as inputs measures the outcome twice. Pick the window from your own cycle, keep only deals still open at the end of it, and run two alternatives as a check.

Our activity is auto-logged. Does that break this?

It changes which metrics are readable, not whether the analysis runs. Auto-logged email inflates volume evenly across won and lost, which is survivable. A sequencer firing on a schedule is worse, because it makes touch counts a property of the cadence rather than the rep. Say so and those rows are excluded.

What use is a metric the buyer controls?

Considerable use, in the right place. Inbound replies separate won from lost more sharply than anything a rep does, so they belong in deal scoring as an early warning. They do not belong on a scorecard, because instructing a rep to receive more replies is instructing them to be handed better accounts.

Is correcting for fourteen tests not overcautious?

The alternative is worse. Fourteen tests at 0.05 yield 0.7 false passes on average, and in the worked example exactly one of the seven that cleared did not survive correction. Every metric tested is listed whether it passed or not, so you can see the denominator behind the shortlist. The same threshold discipline applied to a manager's own forecast history is the forecast submission pack.

Does this prove multithreading causes wins?

No, and the document says so in those words. Deals where a second contact was reachable may simply have been better deals. What the sheet gives you is a ranked hypothesis and the experiment that would settle it, which is where the multithreading plan picks the work up.

Do we have the deal volume for this?

For a large effect, almost certainly. Detecting the 35.6-point gap in the worked example needs 30 opportunities per arm, about six weeks of pipeline. A five-point refinement needs 2,522 and four years. The sheet prints both, and the finding then goes into the deal review.

Sales Activity Analysis Against Win Rates

Fill in the form and your workspace opens with the work already underway.