Research & PolicyFree
Public Comment Docket Analysis Report
Identical submissions are clustered into the campaign that produced them, so a docket of forty thousand resolves into the arguments an agency has to answer.
River clusters the docket by text before it reads anything. Comments that are identical or substantively identical collapse into the campaign that produced them, with the organiser named, the source paragraph quoted and the submission window shown. What is left is the set of submissions containing content nobody else filed. Those get coded one at a time: the position taken, the type of argument made, the category of commenter, and whether the agency is obliged to respond to it.
The count everyone quotes is the wrong number, because an agency does not tally a docket. It has to respond to comments that raise a significant issue, and thirty thousand copies of one letter raise one. The Administrative Conference recommends agencies run de-duplication software that identifies the unique content of comments for exactly this reason. Nothing sold as docket analytics does it; they report sentiment and volume, which is the poll reading again, in a dashboard.
Built for the advocacy analyst working out whether their campaign registered, the counsel checking whether the argument that matters to their client is on the record, and the researcher who has to describe what a docket collectively argued. Run it on the rule the filing analysis already took apart, code it with the discipline of an evidence extraction table, and keep the sources the way a documented search strategy does. More research tools and workspace packs alongside.
Forty-one thousand comments, thirty-four arguments
The docket on that specialty-coating rule closed with 41,284 comments. Clustering by text finds three campaigns: 28,910 identical letters from a consumer group, 9,440 from a trade association's members, and 1,602 from a second consumer campaign in support. Another 703 are form letters with a sentence or two added. That leaves 629 written from scratch. Three campaigns and 629 individual submissions, out of a number the trade press reported as overwhelming opposition on the day it closed.
The 703 edited form letters matter more than they look. 512 add nothing the form paragraph did not already say, but 191 add an operational detail from the writer's own site, and that added text is unique content the agency has to read. With the 629 written from scratch and the three campaign letters themselves, that is 823 unique texts. Coding them by argument type collapses 823 into 34 distinct arguments, which is what the agency's response section actually has to work through.
Six of the 34 are the ones with teeth. Eleven comments argue the agency lacks authority over repackagers. Forty-seven challenge the methodology behind the 41-hour burden estimate, and twelve of those supply their own timings. Those are arguments the agency has to work through, because a rule is adopted only after consideration of the relevant matter presented, and they came from 58 submissions inside a docket of 41,284. The 28,910-strong campaign gets one paragraph, because it made one argument, and made it thirty thousand times.
How it works
Point at the docket
The docket identifier, or the comment files you already downloaded, plus what you need to learn.
Cluster by text
Identical and near-identical submissions collapse into campaigns, each with its source paragraph and window.
Code what remains
Each unique submission gets a position, an argument type, a commenter category and a quote.
Rank the arguments
The distinct arguments, ordered by whether the agency is obliged to answer each one.
What you get
- The deduplication cascade as a Sheet: every comment assigned to a campaign or to itself
- Each campaign fingerprinted: the organiser, the source text, the submission window and the position
- Every unique submission coded by position, argument type and commenter category, with the quote
- Form letters with added text separated out, because the added sentences are unique content
- The count that matters: distinct arguments, ranked by whether the agency has to answer them
- A Doc setting out each side's case in its own strongest form, quoted
Common questions
Isn't the number of comments the point?
No, and no agency treats it that way. A rulemaking is not a referendum, and the obligation is to answer comments raising a significant issue rather than to tally positions. The volume still gets reported, and it still carries political weight. It just is not the thing that changes a sentence in the final rule.
How do you tell a campaign from people who happen to agree?
Three signals together. Text overlap above a stated threshold, a shared source paragraph that traces back to an organiser's published template, and a submission window that clusters around one send. Any one alone is weak. All three, with the numbers shown, is a finding you can put in front of a board and defend.
So a form letter counts for nothing?
It counts as one argument, argued by many people, and the analysis says exactly that with the number attached. Campaigns demonstrate breadth of concern and agencies do report them. What they cannot do is multiply. A campaign of thirty thousand and a campaign of three hundred both make one argument, and the difference is political, not legal.
What about fake or computer-generated comments?
Flagged, never asserted. The patterns that suggest false attribution or automated generation are real and worth surfacing: repeated identity formats, impossible submission rates, text with no human variation. Those go into the Sheet as flags with the evidence behind each one, for a person to judge. The analysis does not accuse anybody.
Can it tell me whether my own comment landed?
Yes, and this is what most people actually want. Your submission is located in the docket, coded like every other, and reported back as either one of the distinct arguments or a duplicate of one somebody else made better. If you ran a campaign, you get its cluster, its window and its single argument.
What if my docket has two hundred comments, not forty thousand?
Then the clustering takes a minute and the coding is where the value sits. A small docket usually means every comment is written from scratch, so the argument register is longer relative to the docket and the commenter categories matter more. The method does not change, only how much of it is deduplication.
What do I get back?
A Sheet with one row per comment, its cluster, its codes and a quote, and a Doc setting out each side's case in its strongest form with the arguments the agency must answer named first. Bring the rule itself to the filing analysis, or the statute behind it to the bill analysis. The coded positions feed a stakeholder position map directly.
Public Comment Docket Analysis Report
Fill in the form and your workspace opens with the work already underway.