River
Y CombinatorBacked by Y Combinator

Writing & MediaFree

FOIA Document Dump Search Index

River indexes every document, threads the correspondence, and tells you how many pages your keyword search could not read at all.

Start here

River reads the release before it searches it. Every page gets graded on whether text can be pulled off it at all, and the pages that fail become a queue with a page count rather than a silence. Then the index: one row per document with its date, sender, recipient, subject and page range, sortable, plus the correspondence threaded back into conversations. The doc names the documents the rest of the dump points at most often, and the records it names that were never produced.

The pages ranking for this term send you to a tool and stop. Upload it, run the OCR, search your keywords, and the printed answer above them says the same. None of them mentions that a keyword search over a scanned release is a search over a lossy transcription, so the pages where the decoder failed are precisely the pages a search cannot answer for. A hit count with no coverage figure beside it reads like completeness.

Built for reporters working a records release, researchers handed an archive as images, and anyone holding thousands of pages with a deadline. Reach for it the day the dump lands, before the first keyword. If the release has not arrived yet, the FOIA request pack asks for the export rather than the images and counts the statutory deadline. When two documents inside it disagree, conflicting report reconciliation works out which figure the records support.

Your keyword search has a coverage number, and nobody reports it

Optical character recognition is a transcription, and it fails on exactly the material a records office produces: photocopies of faxes, stamped pages, handwriting in a margin, a table scanned at an angle. Tesseract's own guidance is blunt about the floor, wanting at least 300 dpi and warning that noise it cannot remove drops accuracy outright. A page that yields fourteen characters is not a searchable page, however confident the decoder is about those fourteen. Your search will never return it, and it looks identical to a page with nothing on it.

Take a real-shaped release: 318 documents and 4,812 pages, 3,228 of them images. Grading each page on mean confidence, with a floor on how much text came out, gives 2,412 clean, 604 marginal, 212 failed. So 83.0 percent of the release is reliably searchable. Searching one vendor's name over those pages returns 41 mentions. A variant search of the marginal pages finds 6 more. Reading the 212 failed pages finds 9 more, in 5 documents nothing had surfaced. The dump holds 56, the search saw 41, and the earliest mention was in the hand-read set.

So report the coverage figure beside the hit count, and turn the failed pages into a queue somebody can finish. Then do the thing no keyword search does: rebuild the correspondence into threads and check each one against itself. In that release, 174 documents were email and threaded into 38 conversations, of which 21 were complete. Twelve were missing a reply the surviving messages quote back, and five named an attachment nobody produced. That is 45 percent of the threads with a hole in them, and 45 named records for a follow-up request.

How it works

  1. Send the release

    Attach the files or describe what arrived, say what you are trying to find out, and choose the reading order.

  2. Grade every page

    River pulls text where it can, grades each page on how much came out, and reports what a search misses.

  3. Read the index

    A sortable Sheet, one row per document, plus a Doc naming what to read first and the records never produced.

  4. Work it in chat

    Ask for the pages mentioning a name, one month's traffic, or the follow-up request drafted from what is missing.

What you get

  • A sortable document index: date, sender, recipient, subject, page range, type, and each document's extraction grade
  • The coverage figure: pages reliably searchable, pages marginal, and pages no keyword search can reach at all
  • A hand-read queue of the failed pages and a fuzzy-search queue of the marginal ones, both counted
  • The correspondence rebuilt into threads, each checked for replies quoted but absent and attachments named but not produced
  • A read-first ranking by how many other documents point at each one, which finds a release's spine
  • Date coverage against the period you asked for, so a run of empty months reads as a gap
  • The records the release names and does not contain, which is the raw material of your next request

Common questions

Tools already exist for this. Why another one?

Purpose-built tooling exists because the problem is universal, and it does the upload and the search well. What it does not report is its own coverage: how many pages the extraction failed on, and therefore how much of the release your search never looked at. This pass produces that number first.

Does it actually read scanned pages, or just index the file names?

It pulls text off the pages, then grades each one on how much came out. Clean pages go into the searchable set, marginal pages go into a variant-search queue, and pages under a character floor go into a hand-read queue with a page count. Nothing is silently counted as empty.

Why is the number of pages that failed extraction worth knowing?

Because the failures are not random. They cluster on faxes, stamps, handwriting and skewed scans, and Tesseract's own guidance explains why: below roughly 300 dpi, and with noise binarisation cannot remove, accuracy drops. Those are the pages a records office produces most.

What does threading the correspondence give me that a search does not?

A conversation instead of a stack of messages, and a way to see what is absent. Surviving messages quote replies that were not produced and name attachments that never arrived. In the worked release, 17 of 38 threads had a hole in them and 45 records were named but not released.

Can it tell me the records I should ask for next?

Yes, and that is often the more valuable half. The list of records the release names and does not contain, plus any run of months with no documents at all, is a specific follow-up request rather than a complaint about effort. The FOIA request pack files it and dates it.

The release came as images. Could I have avoided that?

Usually. An agency must provide records in the form you requested where they are readily reproducible that way, under the form or format provision, and a request for "electronic copies" is answered lawfully by scanned printouts. Name the system, the fields and the file format next time.

What happens after the index, when I am writing?

The documents become sources, and every claim needs to rest on one. The fact check pack joins each assertion in the draft to the source behind it, interview transcript workup handles the tape side, and source tracking records what each person agreed to.

FOIA Document Dump Search Index

Fill in the form and your workspace opens with the work already underway.