Research & PolicyFree
Survey Sampling and Weighting Methods Note
Send the frame, the dispositions and your benchmarks, get a note that names the source and vintage behind every weighting variable.
River starts with the benchmark rather than the weight. For each variable you plan to adjust on, it asks which document the population figure comes from and what vintage that document is. Then whether the figure is a census of your frame or itself an estimate carrying sampling error, and whether the variable is measured the same way in both places. Variables that survive get weighted. Variables that do not get named in the note as unweighted, with the reason. Then it computes what the weighting cost you.
Search this and page one explains what weighting is. Raking, post-stratification, propensity adjustment, a worked example with three age bands and round numbers. Useful for understanding the mechanics and silent on the thing that decides whether a reviewer accepts the weights: where the population targets came from. A weight matched to a stale benchmark, or to a benchmark whose categories do not align with your questionnaire's, is arithmetically perfect and substantively wrong.
For the author writing a methods appendix, the analyst answering a reviewer who asked how the weights were built, and the team that has to reproduce this next wave. It assumes a cleaned file with dispositions coded, which splitting metadata from responses and logging exclusions produces. The effective sample it computes is what a results report carrying each subgroup's base reports against, and it governs the free-text frequencies coded on a named base the same way. A survey repeated every wave still needs its wording locked wave to wave, a separate concern.
Twenty three people carrying a sixth of the estimate
A membership survey. 4,600 drawn, 239 confirmed ineligible, 1,142 complete and 96 partial. The published standard definitions set the minimum rate as the number of complete interviews divided by the number of interviews plus the number of non-interviews plus all cases of unknown eligibility. Every unknown-eligibility case counts against you in that denominator. It comes to 4,361 here, so the rate is 26.19 per cent, and counting partials as respondents gives 28.39. Both numbers go in the note, labelled.
Four variables were candidates for weighting. Region and setting come from the roster, which is a census of the frame, so those targets are exact and dated to the day the roster was pulled. Years in field is on the roster too but missing for 31.2 per cent of records, so it can only carry a partial benchmark. Subspecialty has no population figure anywhere, so it cannot be weighted at all and the note says so. Two of four variables have an exact benchmark, and that sentence is the whole methods appendix in miniature.
Now the cost. Post-stratifying on region by setting produces six cells. Five carry weights between 0.69 and 1.00. The sixth, South community practice, came back with 23 respondents against a benchmark share of 15.9 per cent, so its weight is 7.90 and one respondent there counts as almost eight people. Effective sample size is the squared sum of weights over the sum of squared weights: 573 out of 1,142. The design effect is 1.99, the margin of error at fifty per cent goes from 2.90 to 4.09 points, and weighting cost 569 respondents.
How it works
Audit the benchmarks
Source document, vintage, category alignment and whether the target is exact or estimated.
Code the dispositions
Every case placed, then the response rate computed to a definition you can name.
Build the weights
Only on variables with a defensible target, with the trimming rule stated up front.
Price the weighting
Effective sample, design effect, and the cells where a handful of people carry the estimate.
What you get
- Every weighting variable with its benchmark document, that document's vintage, and its category definitions
- Whether each target is a census of your frame or an estimate carrying its own error
- The variables you cannot weight, named as unweighted rather than quietly dropped
- The response rate computed to a published definition, with the disposition counts behind it
- Effective sample size and design effect, so the margin of error reflects the weights
- The cells whose weights are doing too much work, with the respondent count behind each
Common questions
Why lead with the benchmark rather than the weighting method?
Because the method is rarely what fails. Raking to targets converges regardless. What fails is the target. A federal benchmark is specific: population totals by sex, age, race, and Hispanic origin, controlled to be equal to population estimates by weighting area as of a stated July date. A benchmark named only by institution, with no release and no reference date, is where a reviewer's objection lives.
What if I have no benchmark for a variable I know matters?
Then you cannot weight on it and the note says so explicitly. That is a better outcome than the two common alternatives, which are inventing a plausible target or quietly dropping the variable so no reader knows it was considered. An unweighted variable named in the note is a limitation. An invented target is a finding built on a number nobody can check.
How does it decide whether to trim a weight?
By stating the rule before looking at the results, then showing what the rule costs. Trimming reduces variance and reintroduces the bias the weight existed to remove, so both effects get quantified: the effective sample before and after, and how far the trimmed cell now sits from its target. You choose with two numbers rather than a convention.
Does this cover complex sample designs?
It documents what you describe. Stratification, clustering, unequal selection probabilities and multi-stage designs each contribute to the design effect separately, and the note keeps those contributions apart from the weighting contribution. Conflating them is common and it makes the weighting look responsible for variance that the clustering caused.
Can it check for nonresponse bias?
On whatever your frame carries, which is the only honest comparison available. Where the roster holds a variable for everyone drawn, respondents can be compared against the full sample on it directly. In the worked example the most experienced band is over-represented by 9.9 points and the newest by 8.4 the other way, which is a finding about who answered rather than a guess.
Which response rate definition should I report?
Name whichever one you use and show the disposition counts, so a reader can recompute any of the others. The published standard definitions run from a minimum rate that treats all unknown-eligibility cases as eligible up to rates that estimate eligibility among them. The difference between the lowest and highest is often several points, which is why an unlabelled rate is not a number.
What comes back?
A Sheet with the weighting calculation cell by cell, each variable's benchmark source and vintage in its own column, and the effective sample computed. A Doc writing the sampling approach, the response rate with its definition, the nonresponse comparison against the frame, the weighting and its targets, and what the weights cost in precision.
Survey Sampling and Weighting Methods Note
Fill in the form and your workspace opens with the work already underway.