Research & PolicyFree
Secondary Dataset Assessment for Research
Send the codebook and your research question, get a variable-by-requirement matrix and a verdict on whether the data can answer it.
River reads the codebook and the documentation against your research question, one requirement at a time. Each requirement gets a verdict: present and measured the way you need, present but measured differently, present but too sparse in your subgroup, or absent. Then it chains your restrictions in sequence and reports what survives at each step, down to the cell your question is actually about. The run ends in a stated answer about whether the question is answerable here.
Every page on page one stops at whether the variable exists. The custodians do not. CDC's guidance for its own national survey is explicit. The presentation standards require that an estimated proportion be suppressed if it is based on either an actual sample size smaller than 30 or an effective sample size smaller than 30. Effective sample size divides the raw count by the design effect, so a cell can clear the raw floor and fail anyway.
Written for the researcher choosing between three public datasets, the doctoral student whose proposal names a source nobody has interrogated, and the analyst asked whether a question is feasible at all. The answer feeds a methods section written from the code that ran once the source is settled, and a cleaning log with every decision recorded once it is downloaded. Where the analysis is already done, the results section built from your output block is the next step. A figure carrying that thin cell gets checked by the review that starts from the claim.
A cell of 63 the custodian will not let you publish
A national survey with 38,412 examined participants across two combined cycles. The question needs a specific age band, an exposure, an outcome and one covariate, in one ethnic subgroup. Chain the restrictions in order and the counts fall. There are 6,204 in the age band, then 5,118 with the exposure measured because it only ran on a subsample, then 3,847 with the outcome. Then 2,201 with the covariate present, and 418 in the subgroup. That is 1.09 per cent of where you started.
Four hundred and eighteen still sounds workable. Split it by exposure and it is 63 exposed against 355 unexposed. Sixty-three clears the raw floor of thirty, so a check on record counts passes. Apply the survey's design effect of 2.4 and the effective sample size for that cell is 26.25. The custodian's own standard says an estimate built on it must be suppressed. The unexposed cell is fine at 147.9, which is why an average over the subgroup hides the problem.
That is the verdict, and it arrives with the fix. Pooling a third cycle lifts the exposed cell from 63 to roughly 94, an effective 39.2, which clears the floor. On the requirement side, four of nine requirements are fully met, two are measured differently, two are too sparse and one is absent. The absent one is the exposure as you defined it, so the question is answerable only in a narrowed form. All of that is an afternoon's work rather than a month's.
How it works
Decompose the question
Every variable your question needs, and the form each one has to take.
Read the codebook
Each requirement matched against what the documentation says the dataset actually contains and measures.
Chain the filters
Every restriction applied in sequence, with the surviving count recorded at each step.
Return the verdict
Whether the question is answerable, and the narrower version the data would support.
What you get
- Every requirement from your question scored against the variables that actually exist
- Variables that are present but measured differently, kept separate from ones genuinely absent
- Your restrictions chained in sequence, with what survives reported at every step
- Effective sample size as well as raw counts, because the design effect decides publishability
- The cells the custodian's own standard would suppress, identified before you analyse them
- A stated verdict on answerability, plus the narrower question the data would carry
Common questions
Why does effective sample size matter more than the record count?
Because a complex survey is not a simple random sample. Clustering and weighting inflate the variance, and the design effect quantifies by how much. Dividing the raw count by it gives the sample size your estimate behaves like. A cell of 63 with a design effect of 2.4 behaves like 26, which is below the floor the data's own custodian publishes.
What if the dataset is not a complex survey?
Then the design effect is one and the effective count equals the raw count, which the run states rather than silently skipping. Administrative records, registries and cohort extracts each carry their own hazards instead: linkage failure, differential coverage, and variables that change definition between collection years. The chain still applies and the reasons for each drop differ.
Can I do this without downloading the data?
Mostly, and that is the point. Codebooks, documentation and published frequency tables carry enough to build the requirement matrix and estimate most of the chain. The step that needs the file is the exact joint missingness across your variables, which is usually worse than the marginals suggest. The run says which figures are estimated and which are counted.
Should I just pool more years to get the numbers up?
Sometimes, and the run computes what pooling buys before you commit. The cost is that definitions drift between cycles, so a pooled variable can mean two things. It also assumes the quantity is stable across the pooled period, which is exactly the assumption a trend question cannot make. The note states both costs alongside the gain.
What counts as measured differently rather than present?
A variable that names your construct but operationalises it another way. Self-reported rather than measured, a three-level band rather than a continuous value, a twelve-month recall rather than a point prevalence. Each is usable and each changes what your finding means. The matrix records the difference so it lands in your limitations rather than surprising a reviewer.
What if the verdict is that the question is not answerable?
Then you have saved a month, and the run does not stop there. It states the narrowest version of your question the data would carry, the restriction that cost the most, and what a different source would have to provide. Sometimes a power grid across plausible effects shows the surviving cell was never going to work anyway.
What do I get back?
A Sheet with the requirement matrix, one row per requirement with its verdict and the variable that satisfies it, plus the filter chain with the count and effective count at every step. A Doc stating what the data measures, its coverage and missingness, and the questions it cannot answer.
Secondary Dataset Assessment for Research
Fill in the form and your workspace opens with the work already underway.