How we cleaned 40,000+ HubSpot records without guessing
Deduplication, normalisation and segmentation for LotusConnect's HubSpot CRM, using confidence scoring, a human review queue and a full audit log.
1. The problem
The CRM had grown for years without data standards. The same school district appeared under several names ("NYC DOE", "New York City Dept of Education", "NYC Department of Ed"): three records, three contact lists, no overlap. Here's what was broken:
| Issue | Example | Scale |
|---|---|---|
| Duplicate companies | "NYC DOE" vs "NYC Dept of Education" | 5,813 companies in fuzzy-match groups |
| Missing vertical | Field blank | 40,645 contacts had no vertical |
| Non-standard states | "New York", "new york", "NY", full addresses | 10,643 education contacts fixed |
| Job title sprawl | 1,854 variants of 20 real roles | 7,791 contacts mapped |
2. The method: five buckets, no unclassified records
3. Confidence scoring
Each candidate pair gets a composite score:
4. Order of operations
5. No silent failures
6. Results
7. Why this matters for revenue data
Lead-to-cash breaks in the same places: the same customer under different IDs in HubSpot, Stripe and QuickBooks, fields nobody owns, and writes that fail silently. We apply the same approach to revenue data: deterministic match keys, confidence tiers, a review queue for ambiguous cases, and a log for every change.
Have customers who don't match across your stack?
Book a leakage audit. We'll walk through your lead-to-cash flow and tell you where the joins are breaking.
Book a leakage audit