← News
Automation••7 min read

Automating Data Cleanup: Duplicates and a Single Customer Record

The same company sits in the CRM three times — once with a tax ID, once without, once with a typo in the name — and nobody can say for certain which record is the right one. Why duplicates arise, what can be cleaned automatically, and why merging records without human oversight is often more dangerous than the mess itself.

In most companies that have grown over several years, the same company exists in the system more than once. A salesperson created it once after a phone call, it was created automatically from an online order, and it arrived a third time via invoicing when another company's book of business was taken over. Every record has a slightly different name, a tax ID filled in or not, a different contact. Nobody can say with certainty which one is current.

This isn't a result of carelessness. It's the natural outcome of a customer entering the company through several channels at once, each with its own way of creating a record.

Why duplicates arise

  • Multiple entry channels. Sales, the online store, the invoicing system, a marketing form — each can create a new record if it doesn't notice the customer already exists.
  • Typos and name variants. "ABC Ltd", "ABC Ltd." and "ABC Limited" are three different strings to a system, even though they're the same company.
  • Data changing over time. A company moves address, changes its director, or renames itself after an acquisition — the old record stays and a new one appears alongside it.
  • No unambiguous rule at creation. If the system doesn't enforce a check against a company registration number or another unambiguous identifier before creating a new record, a duplicate is born at the first moment of inattention.
In short: Duplicates aren't one person's mistake — they're a consequence of missing a single entry point. Without one, they reappear at the same pace the company grows.

What can be cleaned automatically

Finding duplicate candidates. Matching by registration number is unambiguous and safe — two companies with the same number are certainly the same company. Matching by name and address is probabilistic: text similarity, a similar address, the same phone number. A system can propose these as a likely pair, not confirm them.

Format normalisation. Standardising how something is written — capitalisation, spacing, diacritics, phone number format. This is safe automation, because it doesn't change the meaning of the data, only its shape.

Filling in missing details from public registers. If a record has a registration number but is missing the legal form or current address, it can be filled in from a publicly available register. We covered this in more depth in our case study on automated verification against public registers; the same principle applies to simpler record enrichment.

Flagging a suspected duplicate at creation. Rather than cleaning up after the fact, it's more effective to prevent a duplicate from being created in the first place — the system warns the salesperson that a similar record already exists before a new one is created.

Why automatic merging without oversight is dangerous

This is where automation most often goes too far. Merging two records means combining order, invoice and communication history into one — and if two companies that merely look similar get merged (the same name in a different city, a franchise partner with its own registration number), the resulting damage is hard to unpick.

SituationAutomateWhy
Matching registration numberyes, merge automaticallyan unambiguous identifier
Name and address match above 90%propose for approvalprobable, not certain
Similar name, different citydon't offer as a duplicatecould be a branch or a different company
Same phone number, different namepropose for approvalcould be the same contact person at a company
Caution: Merging two records is an operation you need to be able to reverse. If two different companies get merged by mistake, their histories mix together — without a backup of the original state, that can only be untangled manually and slowly.

Who is the source of truth

Before anything automated is switched on, it has to be clear which system is the source of truth for which field. If both the CRM and the invoicing system hold a customer address and they differ, an automation with no rule for this doesn't know which one to overwrite and which to keep — it just moves the uncertainty from one place to another.

A typical split:

  • Invoicing and legal data (registration number, tax ID, registered address) — the source of truth is usually the accounting or ERP system, or the public register directly.
  • Sales contact details (person, phone, email) — the source of truth is the CRM, where a salesperson updates them continuously.
  • Preferences and communication history — the source of truth is the system where the communication actually happens.

This split ties into how systems get connected in the first place — we discussed a similar reasoning in our article on connecting ERP, CRM and an online store, where a single source of truth per data type is a precondition, not an optional nicety.

Why it's an ongoing process, not a project

A one-off database cleanup delivers a visible effect immediately, but without changing how records are created, the mess returns at the same pace. In six months the database will be full of duplicates again, unless the process for creating new records changes — a check at entry, a unified form, a mandatory field for the registration number wherever it's available.

An effective approach combines a one-off cleanup of existing data with a permanent change to the process that prevents new duplicates from forming. Without the second part, the first is just temporary cosmetics.

Summary

Duplicates in customer data arise from multiple entry channels, not carelessness, and cleaning them up only makes sense combined with a permanent change to how records are created. A match on an unambiguous identifier like a registration number can be merged automatically; a probabilistic match should only be proposed for approval. And before anything is automated, it has to be clear which system is the source of truth for which field.

The scope of a project like this depends on how many systems hold customer data and how much they differ today. If you're dealing with a similar situation, we'll go through it in a no-obligation consultation.

INTERFASE