CRM Data Hygiene Checklist
Survivorship written field by field before a single pair is reviewed, plus the eleven fields your platform resolves whatever your reviewer clicks.
Free download · No account needed
Every CRM hygiene checklist runs on the same cadence: log activity weekly, deduplicate monthly, audit the fields quarterly, rebuild the schema once a year. It is reasonable advice about reversible work, and it gives one line to the only operation in the whole job that destroys data permanently. A merge cannot be undone. Not by support, not by an API call, not by recreating the record that disappeared. So this pack works outward from the merge instead.
Survivorship gets written field by field before a single pair is looked at, because at merge time a reviewer has ten seconds per pair and keeps whichever record looks fuller. That reflex is how a working phone number gets replaced by one disconnected in 2019. The older create date wins on account age. The more recently verified value wins on phone. The owner of the record carrying the open deal wins on owner, whichever side is primary.
Then the half nobody writes down: eleven fields your platform resolves on its own logic, whichever value your reviewer clicks. Lifecycle stage ratchets to the furthest down the funnel, so a fresh lead merged into an old customer stops being a lead, which is why stage definitions belong next to this. Marketing contact status ratchets to the more marketable of the two, which is a bill rather than a data point. The sheet marks those rows not overridable, because knowing them is the entire value.
What the pack does that a cadence checklist does not
Writes survivorship as rules that resolve
One row per field, naming the deciding attribute rather than the preference. The older create date wins on account age. The legal entity name wins on company name. Two people applying the sheet to the same pair reach the same answer, which is the only test a merge-time rule has to pass.
Names the fields you cannot control
Eleven rows marked not overridable, with the platform's own behaviour quoted and the consequence spelled out. Three of them ratchet: lifecycle position, billable status and consent basis. The output is not a fix, because there is none. It is knowing which numbers move during a cleanup.
Treats the snapshot as the audit trail
The exportable merge history reaches back ninety days and covers only merges run in the duplicates tool, so a one-off merge from a record's actions menu leaves no export. Any field the merge updates is restamped with the merge date. The snapshot is the only copy of the original dates.
Splits the queue by what a criterion can resolve
Exact-email and exact-identifier pairs with trivial associations go through in bulk and get logged. Pairs carrying an open deal, a hand-built list, a hierarchy or a different owner land on the candidate list with a named primary. Five mechanical conditions decide which, not judgement.
Gives blocked pairs a real state
A hierarchy that fails the merge, a combined association count above the limit, and a pair whose records have exhausted the merge-count ceiling are three different problems with three different next actions, one of them unrecoverable. Blocked pairs are the only category that gets worse when ignored.
Separates fill rate from correctness
Fields that default on creation are always fully populated and frequently wrong, so a hygiene score built on fill rate reports a bad database as healthy. The compliance sheet counts blanks on records with an open deal separately, and marks the fields where full compliance still means bad data.
How the pack runs
- 1
Send the export
Contact and company export with record id, create date, owner, email or domain, and whatever fields your reports group by. A current duplicate pair export is welcome alongside it.
- 2
Write survivorship first
River fills the field-by-field sheet with your rule, the platform default beside it, and whether the value can be overridden at merge time at all.
- 3
Measure the detection rules
Each rule gets a judged sample and a false positive rate, which decides whether it can be merged from in bulk or is triage only.
- 4
Review before anything merges
The candidate list arrives with a named primary per pair, the conflicts listed, what disappears written down, and a decision of merge, reject, hold or blocked.
Frequently asked questions
Deduplication is one line on a hygiene checklist. Why a whole pack?
Because it is the only line that cannot be undone. Fill rates, formats and stale records are all recoverable next quarter. A merged record is gone, and the surviving one now carries a mixture of two records' values that nobody chose deliberately. Everything else on the checklist can wait.
What does it mean that a field is not overridable?
The comparison dialog highlights your choice in green and the platform resolves the field by its own rule anyway. HubSpot documents that lifecycle stage keeps the stage furthest down the funnel and marketing contact status keeps the most marketable of the two. Both are ratchets, and one of them is billable.
Can we just bulk merge and fix the problems afterwards?
For exact-email pairs with trivial associations, yes, and the pack says so. The five bulk survivor criteria are most recent engagement, oldest engagement, created first, created last and most recently updated. None is the record with the open deal attached, so anything carrying live pipeline needs a human.
There is no undo, but is there not a merge history to fall back on?
Thinner than it looks. The exportable history reaches back ninety days and covers only merges performed in the duplicates tool, so a one-off merge from a record's own actions menu leaves nothing. And updated fields are restamped with the merge date, so field history loses the original dates too.
How do we know which fields even need a survivorship rule?
From what your reports group by, not from the field list. A field nothing consumes does not need a rule and probably should not exist. If you have not mapped that yet, a CRM field dictionary names every field, its owner and the report consuming it, which is the input this sheet is built from.
Our duplicate backlog is bigger than the tool will display. Now what?
Count it from the export instead. The duplicates manager displays a bounded number of pairs by subscription tier, so above the ceiling the tool cannot show you the size of your own problem. Sizing the whole cleanup as a priceable scope is what a CRM audit diagnostic does.
We run two CRMs and both are full of duplicates. Where do we start?
Inside one system, with survivorship, because a merge in either one is irreversible while a sync disagreement is not. Once each side is clean, a HubSpot and Salesforce reconciliation classifies the remaining record-level differences by cause instead of merging blind.
Write survivorship before the next merge
Send the contact and company export and get the field-by-field rules, the detection rates and the candidate queue back.
Get the template