River
Y CombinatorBacked by Y Combinator

Non-ProfitFree

Donor Database Duplicate Cleanup Tool

Send the CRM export, get every probable duplicate group with a proposed survivor field by field, and the form that created it.

Start here

River's Donor Database Duplicate Cleanup Tool reads your CRM export and returns one row per probable duplicate group, each carrying the intake path that created it and a survivor assembled field by field rather than chosen whole. Nothing is merged. The sheet is a decision queue for a person to work through. Beside it sits the entry standard those same groups imply, because merging fixes the records you hold today and changes nothing about the form still making more of them tomorrow.

Every cleanup guide gives you the same routine. Run the duplicate report, sort by name and email, review each match before merging, repeat monthly. That treats a duplicate as untidiness. A duplicate is really a receipt for one form failing one match check on one day, and the receipt is still sitting in the file. Attributing groups back to the intake paths turns an export into a ranked list of rule changes, which is what segmentation off real giving history needs before it can be trusted.

This is for the development director who inherited a file nobody trusts, the operations manager mid-migration who does not want the duplicates carried across, and whoever just fielded a complaint about two letters in one week. A gift that was never recorded at all is a different failure, and three-way reconciliation is what finds those. A gift recorded against the wrong twin still reaches the acknowledgment queue, addressed to a version of the donor's name they never use.

What a split giving history actually costs

Start with the junk records, every guide says, and it points at the wrong end of the file. A matching rule only gets a chance to fail when somebody gives, so duplicate rates climb with giving frequency and the most committed donors are the most exposed. The stock Salesforce NPSP rule matches on fuzzy first name, exact surname and exact personal email, so one gift made from a work address is enough to open a second record.

Ashgrove Literacy Partnership exported 18,799 constituent records covering 81,885 gifts. Grouping found 1,402 probable duplicate groups holding 1,559 surplus records, 8.3% of the file. In the top lifetime-giving decile 30.0% of people carry a duplicate against 1.9% in the bottom, and the median top-decile donor made fourteen separate gifts against two. One event import that ran no match check produced 41% of the duplicates on 11% of the transactions, while the two paths staff key by hand produced 6% on 28%.

Recombining the groups moved $3,131,124.39 of giving onto records somebody reads. Forty-eight donors have given $170,729.55 between them and not one appears on a $2,500 major-gift review list, because no single record of theirs reaches it. Another 543 read as lapsed or single-year donors on every record they hold, while the person behind them gave in three or more separate years. So the ask gets derived off half a history and then sent twice, and the median understatement of a donor's own previous best is $49.05.

How it works

  1. Send the export

    Every CSV in the zip, constituents and gifts both, whatever your CRM decided to call the columns.

  2. Group, then attribute

    Probable duplicates grouped, then each group traced back to the form or import that created the extra record.

  3. Price what it hides

    Recombined giving per person, who a merge moves onto a review list, and who only reads as lapsed.

  4. Decide, then prevent

    One merge decision per group for a person to sign, then the rule change each intake path needs.

What you get

  • Every duplicate group traced to the intake path that created it, with that path's failure rate
  • A survivor built field by field, because the record with the good address holds the dead email
  • Suppression flags unioned across the group, so no merge can quietly re-enrol somebody who opted out
  • Recombined lifetime giving per person, and who crosses your major-gift threshold only once merged
  • Match candidates that must not merge, pulled out separately, with couples and namesakes named as such
  • The entry standard as one specific rule change per form, ranked by the duplicates each would prevent

Common questions

Will it merge the records for me?

No, and that is deliberate. A merge in a real CRM overwrites every field you did not pick and cannot be reversed. So the output is a decision queue: one row per group, a proposed survivor per field, and the groups where a wrong pick would drop a do-not-solicit flag sorted to the top. A person signs each one.

How is this different from my CRM's duplicate report?

A duplicate report hands you match candidates. It does not say which form made them, and that is the only part that stops the next batch arriving. In the worked example one event import ran no match check at all and produced 41% of the duplicates on 11% of the transactions, so it is one fix rather than 1,402.

Isn't the real answer a monthly dedupe routine?

A routine holds the count flat. It never lowers it. Ashgrove's five worst forms account for 94% of the duplicates in the file and generate roughly 260 new ones a year, so merging without changing the rules starts decaying the day it finishes. Keep the routine. It is not the fix.

What about couples who share an email address?

They are the largest false positive in the file, so they get their own list rather than a merge row. Same surname and address with a different first name returned 399 candidates that must not merge, 214 of them couples, and merging one loses a constituent you are separately in touch with. Namesake parents and children behave the same way.

Which record should survive?

Usually none of them whole. The record with the verified address often carries the dead email, and the record a gift form made holds the name the donor uses while the staff-keyed one holds the name on the cheque. So survivorship is decided per field, and the earliest first-gift date in the group always wins.

Does this actually stop the duplicate letters?

It prices them first. Ashgrove's 1,504 mailable surplus records cost $7,459.84 a year across four mailings at $1.24 a piece. A duplicate is also the address nobody updates, and the postal Move Update standard wants every address refreshed within 95 days of the mailing date, against a 0.5% error threshold.

Does it read Raiser's Edge and Bloomerang exports?

Yes, including the multi-CSV zip. Column names differ between systems and between two exports from the same system, so the mapping is stated back to you before anything is grouped. The same file then settles what the board is told about donor counts, which 1,559 surplus records overstate.

Donor Database Duplicate Cleanup Tool

Fill in the form and your workspace opens with the work already underway.