River
Y CombinatorBacked by Y Combinator

Business & Revenue OpsFree

HubSpot and Salesforce Data Reconciliation

Record-level reconciliation of both exports, with every difference classified by what caused it and the count each cause explains.

Start here

Both numbers are defensible and only one of them is going in the deck. The sync error log is the first place everybody looks and it is almost never where the gap lives. The three largest divergences between a HubSpot deal export and a Salesforce opportunity export are documented behaviours of the connector rather than failures of it, and a skipped record raises nothing at all. It is simply absent on one side, in a report nobody filtered the same way.

Start with the filters, because there are two of them running in opposite directions. HubSpot's own documentation on inclusion segments states that records outside the segment will not sync, and selective sync applies a separate filter to everything coming back the other way. They live on different settings screens. Neither one is wrong, neither one logs anything, and between them they decide which records were ever eligible to appear in both systems at all.

Then the ratchet. HubSpot documents that the default lifecycle stage property can only be moved forward by its tools, the Salesforce integration among them. A deal lost in Salesforce therefore leaves its contact at Customer until a person clears it by hand. Add that Opportunity is a stage on a contact in one system and an object in the other, and two teams can count opportunities correctly and disagree by hundreds. What comes back is the reconciliation at record level, every difference typed by cause, with the count each cause explains.

Classify by cause, not by row

Once the filters are accounted for, the remaining gap splits along the object model. A Salesforce opportunity hangs off exactly one account. A HubSpot deal can be associated with several companies at once, so an account-level rollup counts the same deal more than once on one side and once on the other. Nothing failed. The two systems answered two different questions and both answers are internally consistent, which is why the meeting goes in circles.

Sync failures are a real category and they are usually the smallest one. They are also the only category anybody has tooling for, which is why the audit that gets run is an error audit and the question that gets asked is a counting question. Ordering matters here: work the filters first, then the definitions, then the associations, then the failures, then whatever is left. Anything you fail to attribute at each step is not evidence of a bug, it is the input to the next step.

The residual is what you actually wanted. Records present in one system and genuinely absent from the other, with nothing structural explaining them, are the list somebody has to work. Everything above them is a definition problem, and definitions are fixed once rather than reconciled monthly. That is why the output ends by naming which system is authoritative for each field in dispute, in the form a field dictionary can carry forward. Where the residual is large enough to be a project, it becomes a line in a costed cleanup scope.

How it works

  1. Send both exports

    The HubSpot deal or contact export and the Salesforce report behind the other number.

  2. Match the records

    Pairing runs on whichever identifier survived the sync, then on email, domain and amount.

  3. Attribute the gap

    Each difference gets a cause, worked in order from filters through definitions to genuine absence.

  4. Answer the question

    Which number to quote today, why, and the single definition change that ends the argument.

What you get

  • Record-level reconciliation of both exports, matched on the identifier each system actually keeps
  • Every difference classified as a filter, a definition, an association, a sync failure or a real gap
  • The count of records each cause explains, so the total gap resolves to zero
  • Records the inclusion segment or selective sync filtered out, which never appear as errors
  • Contacts stuck at a lifecycle stage the integration is not allowed to move backwards
  • Which system is right per disputed field, and the one change that stops the recurrence

Common questions

What do I send?

Both exports. From HubSpot, the deal or contact export behind your number. From Salesforce, the report behind theirs, including whatever filters it uses, because the filter is the single most common cause and it is invisible in the exported rows. Screenshots of the two report definitions are enough if you cannot export them.

Our sync error log is clean. Doesn't that mean the data agrees?

No, and that is the assumption this exists to break. A record excluded by an inclusion segment or a selective sync filter is skipped rather than failed, so it produces no error and no notification. The largest gaps are usually made entirely of records that were never eligible to cross, which the error log has no reason to mention.

Which number is actually right?

Whichever one answers the question being asked, and the output says which for each. Total open pipeline is usually the HubSpot figure because it includes pipelines the connector never mapped. Anything a forecast depends on is usually the Salesforce figure because that is where the stage discipline lives. The point is deciding once, in writing.

Why is our customer count higher in HubSpot every quarter?

Because the lifecycle stage only moves one way. HubSpot documents that the default property can be moved forward by its tools, the Salesforce integration included, and moving it backwards takes a manual clear or a workflow. Every deal that closes lost after a contact reached Customer leaves that contact at Customer permanently. Correcting it needs stage definitions with buyer-evidenced exits, not a bulk update.

Does this apply to any two systems, or just these two?

The method is general and the specifics here are not. Two CRMs, a CRM and a billing system, an ad platform and an analytics tool: all of them diverge by filter, definition, structure and genuine absence. For a pair outside this one, reconciling two conflicting reports runs the same classification, and the rest of the operations tool set covers the neighbouring jobs.

How do we stop having this argument every month?

By settling which system is authoritative for each disputed field and what each metric computes, rather than reconciling again. The reconciliation names the fields in dispute and which side wins each one. A register of what your reports actually compute holds those answers permanently, so next quarter the question arrives with a documented answer attached instead of two exports.

HubSpot and Salesforce Data Reconciliation

Fill in the form and your workspace opens with the work already underway.