River
Y CombinatorBacked by Y Combinator

Marketing & GrowthFree

SEO Forecast Model and Business Case

The forecast built on your property's own measured click-through by position, every input labelled, and the whole thing priced against paid.

Start here

River's SEO forecast derives the click-through curve from your own history instead of importing one, then builds the projection on top of it. Every input in the model arrives labelled: measured from your files with the row count behind it, stated by you as a business number you own, or assumed by the model and flagged as such. Each assumption also carries the value at which the case stops working. You get the model as a spreadsheet, the business case as a document, and a short deck.

The templates that rank for this search all have the same hole in the middle. They multiply keyword volume by a published position-to-clicks curve, and that curve is doing most of the work while being the one number nobody sourced. Rundale Systems' own measured curve returns 1,112 sessions a month on its target query set. The same query set on a widely circulated published curve returns 1,591. Nothing else changed, and the forecast grew by 43 percent.

Built for in-house SEOs writing a case for headcount, agency leads justifying a retainer, and anyone whose forecast is about to be read next to a paid media plan. Take the de-duplicated query set from the keyword cluster map, and the full history from the Search Console keyword map rather than the truncated interface export. Where the plan depends on already-ranking pages moving, the internal linking audit says which of them a link could plausibly shift. It runs in the marketing workspace beside all three.

Three inputs, and only one of them is a measurement

Start with position, because it is not a rank. Search Console defines it as the average position of the topmost result from your site, averaged over every impression in the window. So a page reported at 11.4 is not sitting eleventh; it is a distribution with a mean of 11.4, and moving that mean to 5 is a different claim from moving a rank. A forecast that treats the figure as a rank and applies a rank-based curve to it has already double-counted.

Then volume, which is not a count of searches either. Google defines average monthly searches as the figure for a keyword and its close variants, averaged over twelve months and rounded. Rounding alone means a cluster's rows do not sum to what you expect, and the twelve-month average has the seasonality already baked in, so applying a seasonal factor on top counts it twice. The forecast states which of the two it did and never does both.

Which leaves the curve, and it is the only one you can actually measure. Your property already reports its own clicks, impressions and position, so its own click-through by position band is arithmetic rather than a borrowed constant. Split branded from non-branded before you do it: Rundale's branded queries convert impressions at 46.2 percent in the top band against 18.4 percent non-branded, and a published curve is a blend of both. The whole history is available through the API at 25,000 rows a request against the interface's 1,000.

How it works

  1. Send the history

    Query and page rows over as long a window as you have, plus the query set you want.

  2. River measures the curve

    Your own click-through rate by position band, with branded and non-branded kept apart throughout.

  3. Read the labelled model

    Each input marked measured, stated or assumed, with its breaking value beside it.

  4. Argue with the inputs

    Change a stated number, watch the case move, and take the version you can defend.

What you get

  • A click-through curve by position band, derived from your own property rather than a study
  • Branded and non-branded queries kept apart, because blending them overstates every non-branded forecast
  • Every input labelled measured, stated or assumed, with the row count behind the measured ones
  • The value at which each assumption breaks the case, so the argument is bounded
  • A do-nothing scenario, because existing pages decay and the counterfactual is not flat
  • Cost per incremental session against the cost of buying the same clicks, year by year

Common questions

How much history do I need for the curve to be usable?

Enough rows per position band that the band's rate is stable, which for most properties means several months rather than several weeks. The output states the row count behind every band and marks any band too thin to use. A thin band falls back to a labelled default rather than quietly borrowing from the band next door.

What if I have no history at all?

Then the curve is an assumption and the model says so on the face of it. A new property gets a labelled default curve and an explicit note that the single largest input is unmeasured. That is a weaker case, and presenting it as a measured one is how forecasts get discredited on the second question.

Why separate branded from non-branded?

Because they behave nothing alike and a published curve blends them. Rundale's branded queries click at 46.2 percent in the top band against 18.4 percent non-branded, which is 2.5 times. A forecast for pages you have not built yet is a non-branded forecast, so branded rows are excluded from the curve and reported separately.

Is a single traffic number not what the board wants?

They want one number they can defend when someone asks where it came from. The deck leads with a single figure and the model behind it labels every input as measured, stated or assumed, with the value at which each one breaks the case. The scenarios exist so the answer to what if you are wrong is already written.

Why compare against paid at all?

Because whoever approves this will do it anyway, usually in the meeting. Pricing the same volume at the keyword export's own top-of-page bid puts both options in the same units. In the worked example year one costs 2.41 times what buying those clicks costs, and saying that first is what makes the year-two figure credible.

Should I forecast revenue or traffic?

Both, in that order of scepticism. Sessions come from measured and assumed inputs; revenue adds two conversion rates and a contract value that are yours to state. The model keeps them in separate blocks so a reader who trusts the traffic and not the revenue can stop reading at the session line and still have a usable number.

SEO Forecast Model and Business Case

Fill in the form and your workspace opens with the work already underway.