How we cleaned 40,000+ HubSpot records without guessing

Deduplication, normalisation and segmentation for LotusConnect's HubSpot CRM, using confidence scoring, a human review queue and a full audit log.

Client
LotusConnect LLC
Platform
HubSpot CRM
Records
40,647 contacts · 21,653 companies
Outcome
121 active segments
40,647
Contacts processed
923
Duplicate company groups auto-merged
351
Ambiguous groups routed to human review
121
Active segments deployed

1. The problem

The CRM had grown for years without data standards. The same school district appeared under several names ("NYC DOE", "New York City Dept of Education", "NYC Department of Ed"): three records, three contact lists, no overlap. Here's what was broken:

IssueExampleScale
Duplicate companies"NYC DOE" vs "NYC Dept of Education"5,813 companies in fuzzy-match groups
Missing verticalField blank40,645 contacts had no vertical
Non-standard states"New York", "new york", "NY", full addresses10,643 education contacts fixed
Job title sprawl1,854 variants of 20 real roles7,791 contacts mapped

2. The method: five buckets, no unclassified records

1. True singletons
No action
2. Exact duplicates
One master kept, others archived
3. Keyword-resolved
Suffix differences (LLC, Inc, School District) resolved by rules
4. Fuzzy auto-merge
High-confidence matches merged through the HubSpot Merge API (923 groups)
5. Needs review
Ambiguous pairs exported for client sign-off, with no automated action (351 groups)

3. Confidence scoring

Each candidate pair gets a composite score:

Name similarity
Token-set + partial ratio, after stripping legal suffixes
50%
POI type match
Google Places establishment type (e.g. school vs hospital)
20%
Geographic proximity
Google Distance Matrix; different states = zero
30%
Example: Two "Lincoln Elementary" records
In different states: Name match 94 → Score 66 → Routes to human review
In the same city: Name match 94 → Score 91 → Merges automatically

4. Order of operations

1
Classify companies
2
Deduplicate
3
Backfill contact fields from the master company
4
Map job titles to 20 canonical roles
5
Build segments
Key: Merging before backfilling prevents stale data from absorbed records overwriting clean values.

5. No silent failures

HubSpot forward references
797 "forward reference" errors (records already merged into another). We traced each chain to its surviving master and confirmed all 797 resolved instead of treating them as failures.
Unresolvable values
Left blank, not guessed. 8,258 contacts without a company match were excluded from backfill.
Category reclassification
27 "University/College" companies tagged as Education were actually Healthcare. They were reclassified, and 142 contacts had education-only job titles cleared.
Audit log
Every batch write was logged with record IDs, the value written and the outcome.

6. Results

10,222
Companies classified (100% of companies with contacts)
23,381
Contacts with a vertical
10,643
States normalised to USPS codes
7,791
Job titles canonicalised
923
Company groups merged
121
Active segments deployed (3 lifecycle, 4 customers-by-industry, 114 education geography)

7. Why this matters for revenue data

Lead-to-cash breaks in the same places: the same customer under different IDs in HubSpot, Stripe and QuickBooks, fields nobody owns, and writes that fail silently. We apply the same approach to revenue data: deterministic match keys, confidence tiers, a review queue for ambiguous cases, and a log for every change.

Have customers who don't match across your stack?

Book a leakage audit. We'll walk through your lead-to-cash flow and tell you where the joins are breaking.

Book a leakage audit