River
Y CombinatorBacked by Y Combinator

Research & PolicyFree

Survey Results Report and Base Check

Every subgroup carries its unweighted base and its effective base after weighting, so the report only makes the claims those bases can carry.

Start here

River builds the crosstab first, then writes the report from it. Every question by every subgroup you name, with the unweighted base in the cell, the design effect the weights carry, and the effective base after weighting. The margin of error is computed on that effective base, not on the row count. Then each sentence in the findings is checked against the base under it, and the ones the data cannot carry come back as the reason rather than as prose.

A report template cannot know any of this. It has a slot for a key finding per segment, which is an invitation to write one whether the segment supports it or not. AAPOR's disclosure standards ask a probability survey to report its sampling error and say whether it was adjusted for the design effect due to weighting. Weighting a sample to benchmarks with a named source raises the variance of every estimate it touches. The margin quoted from the raw row count is the narrower of the two whenever the weights vary.

Built for the analyst whose fieldwork closed on Friday and whose stakeholders want a deck on Monday, and for anyone who has watched a claim resting on eleven respondents survive three rounds of review. Run it on the file that came out of survey data preparation, against the subgroup tables the instrument pack planned before fielding. The methods writer then documents what this reported, the cleaning record holds the exclusions these bases rest on, and more research packs sit alongside.

The 198 respondents who are really 82

Take a workforce study that closed with 3,412 analysable responses, weighted to an 18,400-person frame. The whole-sample margin of error looks like ±1.7 points. Compute the design effect the weights carry, 1.52, and the effective base is 2,245, so the real margin is ±2.1. Inside the crosstab the gap is wider. Bank and agency staff number 198 in the file, but they carry the heaviest weights, a design effect of 2.41, and an effective base of 82, which is a margin of ±10.8 points.

The draft report said bank and agency staff report unpaid overtime at 61 per cent against 54 per cent for permanent full-time staff. That 7-point gap needs a margin on the difference, not on either estimate, and on effective bases of 82 and 1,438 that margin is ±10.9 points. OMB's survey standards are explicit about the order of operations: make the comparison test before the statement goes into the product. Run it second and you are defending a sentence you have already written.

Then the cell nobody should have run. Carer strain by employer was forecast before fielding at ten respondents a cell, and the routing plan said to report it pooled. The draft reported one employer anyway: 8 of 11 respondents, printed as 73 per cent, effective base 6. The exact 95 per cent interval on 8 of 11 runs from 39 to 94 per cent. OMB's guideline on release criteria asks for the threshold to be set before the estimates exist, which is the only point at which nobody is attached to the finding.

How it works

  1. Send the data

    The cleaned file with its weight column, plus the subgroups the report has to break out.

  2. Build the crosstab

    Every question by every segment, with the unweighted base, the design effect and the effective base.

  3. Set the floor

    You agree the base below which nothing is reported, before any estimate has been looked at.

  4. Write the findings

    Each claim tested against its own base, with the suppressed cells listed rather than dropped.

What you get

  • The crosstab as a Sheet: every question by every segment, with the unweighted base in the cell
  • The design effect computed from your own weight column, and the effective base it implies
  • A margin of error on every estimate, taken from the effective base rather than the row count
  • Subgroup comparisons tested on the difference, so a gap inside its own margin never becomes a finding
  • The findings Doc, with every claim carrying its base and its interval in the sentence
  • Charts that plot the interval, not just the point, and label the base under each bar

Common questions

What if my sample is not a probability sample?

Then it does not print a margin of error as though it were one. An opt-in panel has no sampling frame to draw the interval from, so the crosstab still carries every base and the report says what the bases are and what the weighting was fitted to. Any precision measure gets named as model-based, with the model stated.

What base is too small to report?

You set it, and you set it before you have seen the estimates. The run proposes a floor with a basis, usually an unweighted base plus the effective base after weighting, and asks you to agree the number. Below the floor a cell still appears in the Sheet with its base, and no sentence is written from it.

Why is the margin of error bigger than the one my software printed?

Because most packages print the interval for a simple random sample of that size, and a weighted sample is not one. The design effect is computed from the spread of your weights, and dividing the row count by it gives the effective base. In the worked example 198 bank and agency responses carry an effective base of 82.

Can it just tell me whether two segments are different?

Yes, and it tests the difference rather than eyeballing two intervals. Two estimates whose intervals overlap can still differ, and two that look far apart on a chart often do not. The test uses both effective bases, so the 7-point gap in the worked example sits inside a margin of 10.9 points.

My stakeholders want a number for every segment. What do I give them?

The Sheet, which has every cell with its base, and a report that only makes claims the bases carry. Where a segment is too thin the run offers the two repairs that work: pool it with a neighbouring segment and say so, or report the pooled figure the design already planned for.

Does it handle change since the last wave?

Yes, and the interval belongs to the change rather than to either wave. In the worked example wellbeing reads 3.30 against 3.34 last wave, a fall of 0.04, and the margin on that difference is 0.05. So the reportable finding is no detectable change, which is a different sentence from a fall.

What do I actually get back?

A Sheet holding the crosstab with a base in every cell, a Doc holding the findings with each claim's base and interval in its own sentence, and charts that show the interval. The methods writer picks the whole thing up when the paper needs a methods section.

Survey Results Report and Base Check

Fill in the form and your workspace opens with the work already underway.