River
Y CombinatorBacked by Y Combinator

Writing & MediaFree

Name and Date Extraction From Documents

Send the document set, and get every name, date and amount as a row with its source page, plus the error rate from a verification sample.

Start here

River turns a document set into rows you can sort, and does two things a plain extraction does not. Every value carries the document and the page it came from, so any figure can be rechecked in one click rather than rediscovered from scratch weeks later. And a sample of the extracted values gets verified against the source pages by hand, stratified by whether the page had a real text layer, so the output arrives with a measured error rate instead of an implied one.

Everything ranking for this query is an extraction engine. Upload the PDFs, choose the entity types, receive a spreadsheet of names and dates. The extraction is usually decent and the spreadsheet is usually uncitable, because a value with no page behind it cannot be verified and a set with no error rate cannot be published. Those are two separate failures and every incumbent has both, which is why the output looks finished and cannot be used.

Built for investigative and data reporters working a released document set, and for researchers building a dataset out of paper. Run it after the set is indexed rather than instead of indexing it. Document dump triage makes an image-only set searchable first, dataset cleaning and methodology preserves the original and logs every transformation, and public records dataset request asks the agency for the underlying database when one turns out to exist.

A 5.5 percent error rate is two very different numbers averaged together

Take an invented release of 4,180 pages, built to a real shape. Only 1,290 of them carry a native text layer; the other 2,890 are scans, so 69 percent of the set has to be read optically. Extraction across the whole thing returns 2,314 values, of which 1,402 come from the text-layer pages and 912 from the images. A sample of 200 values, drawn in proportion to those two strata, gets checked against the source pages by hand.

Eleven of the 200 are wrong, which reads as a 5.5 percent error rate, and the Wilson interval on that sample runs from 3.1 to 9.6 percent. Split by stratum it is a different story: 2 errors in the 121 text-layer values, a rate of 1.7 percent, against 9 in the 79 image-derived values, a rate of 11.4 percent. The image half is not slightly worse, it is nearly seven times worse, and the single headline figure hides that entirely.

That split is what makes the set usable. Weighted properly the overall rate is 5.5 percent, so roughly 127 of the 2,314 values are wrong somewhere. The 1,402 text-layer values are publishable with a note; the 912 image-derived ones need a human pass first. Seven of the 11 errors were digit transpositions in amounts, three were misread names, and one was a date, so the risk is concentrated exactly where a reader would care most.

How it works

  1. Send the set

    Attach the documents and say what you need pulled out and what you will publish.

  2. Get the rows

    Every name, date and amount as a sortable row with the page it came from attached.

  3. Read the error rate

    A hand-checked sample split by page type, with the rate stated for each stratum.

  4. Work it in chat

    Ask which values to check first, or to reconcile a name that appears three ways.

What you get

  • One row per extracted value, carrying the document, the page and the surrounding line
  • A measured error rate from a hand-checked sample, not a vendor accuracy claim
  • The rate split by whether the page had a text layer, because the two differ several fold
  • Errors classified by kind, so you know whether amounts or names are the weak point
  • The subset publishable as it stands, separated from the subset needing a human pass
  • Name variants reconciled, with the raw form kept alongside the canonical one
  • Recovery per page by stratum, which is where an under-extracted section shows up

Common questions

Why does every value need a page citation?

Because the extraction is a copy and the page is the original. Evidence law works the same way: under Rule 1003 a duplicate stands in for the original unless a genuine question is raised about authenticity. A cell with no page behind it cannot answer that question and so cannot be defended.

How is the error rate actually measured?

By hand, on a sample drawn in proportion to the strata that matter. In the worked set that meant 200 values, 121 from pages with a text layer and 79 from scans, each one compared against its source page. Eleven were wrong, and the Wilson interval on that sample runs from 3.1 to 9.6 percent.

Why split the rate by page type instead of reporting one number?

Because averaging hides the thing you need. The worked set came back at 1.7 percent on text-layer pages and 11.4 percent on scans, so the single 5.5 percent figure describes neither half. Knowing which stratum a value came from is what tells you whether to publish it or check it.

Can I skip the sample if the tool seems accurate?

Not if you intend to publish the numbers. An unmeasured error rate is not a low one, and a vendor accuracy claim is measured on clean documents rather than on your fax loop and your 1990s scans. The sample is a few hours of work and it is the difference between a dataset and a guess.

What about values the extraction missed entirely?

Recovery is reported per page by stratum for exactly that reason. Records practice treats the gap as a first-class finding: NARA's digitization rules require an inventory identifying gaps in coverage or missing records, plus a documented process for finding them. A section returning three times fewer values per page is usually under-extracted rather than sparse.

The same company is spelled three different ways. Does that break it?

It gets reconciled rather than silently merged. Each row keeps the raw form exactly as it appears on the page plus a canonical form, so a count by company works while the citation still matches what a reader would see. Merges are listed so you can reject any that are wrong.

Is this a substitute for reading the documents?

No, and it is most useful as the thing that tells you which ones to read. Sorting 2,314 cited values by amount or by recurrence surfaces the twenty pages worth an afternoon, which is a different job from the indexing pass that made the set searchable in the first place.

Name and Date Extraction From Documents

Fill in the form and your workspace opens with the work already underway.