What customer deduplication is

Deduplication reduces several records about one person to one profile — a
single customer view. The task sounds technical, but its
consequences show up in reporting: without deduplication a business does not know how many customers
it has and systematically undervalues them.

The problem does not come from bad architecture. It comes from normal user behaviour. The same
person enters the database repeatedly, from different devices, channels and touchpoints, and each
time the system is entitled to treat them as new.

Where duplicates come from

Source of the duplicate How it looks in the data Frequency
Guest → signed in An anonymous profile with browsing history plus a profile with orders Constant
Different devices Phone and desktop as two independent visitors Constant
Several email addresses A personal address for the newsletter, a work one for the order Frequent
Typos and spelling variants Whitespace, case, lookalike domains Frequent
Phone formats With and without a country code, brackets, dashes Frequent
Offline loyalty card A separate record in the loyalty system with no link to the online profile In omnichannel chains
Ordering on someone else’s behalf An order in a relative’s name on the same phone number Regular

The reverse situation is rarer and more expensive: a shared family laptop or a work computer where
one device hides several people. Any merge by device joins unrelated histories.

Deterministic and probabilistic matching

Property Deterministic Probabilistic
Based on Exact match of email, phone, customer ID, card number Device, IP, delivery address, behavioural traits
Precision High Depends on the threshold; errors are unavoidable
Coverage Only records carrying an identifier Wider, including anonymous traffic
Main risk Under-merging — one person stays several profiles Over-merging — two people become one
Where it applies Orders, sign-in, loyalty programme Linking an anonymous session to a known profile

The base link is performed by identity resolution: an anonymous
identifier in the browser or app is matched to a profile at sign-in or checkout, and the previous
history moves into the merged profile. Probabilistic rules on top extend coverage; they do not
replace deterministic ones.

⚠

Probabilistic device merging must not drive operations with sensitive consequences: showing order
history, personal discounts or personal data. An error rate acceptable for product recommendations
is unacceptable where a shopper would see someone else’s information.

Merge rules

Finding the duplicates is half the work. The other half is deciding what the merged profile becomes.
The rules are set in advance, otherwise the result depends on processing order.

Field type Priority rule
Contact details (email, phone) The confirmed value wins; unconfirmed ones are kept as alternates
Name, delivery address The value from the most recent completed order wins
Order and event history Merged in full, with no records lost
Communication consents The strictest wins: a withdrawal overrides a consent
Tracking restrictions Propagate to the whole merged profile
Technical identifiers All retained, with one marked primary

Reversibility is recorded separately. A wrong merge will have to be split, so the profile must keep
the original identifiers and a record of which rows it was assembled from. A merge that discards its
sources turns any error into a permanent one.

Consent is the one category where the strict rule beats convenience. A formally more recent record
may carry consent, but if any merged record holds a withdrawal, the withdrawal takes priority: the
cost of that error is legal, not marketing.

How deduplication errors break metrics

An inflated customer count. Every unmerged duplicate is a new customer in the report. The share
of new buyers looks healthy while there is no real inflow.

Understated LTV and purchase frequency. Five orders from one person spread across three profiles
read as three customers with one or two purchases each. LTV is understated,
retention looks worse than it is, and acquisition payback is computed wrongly.

Broken attribution. Saw the ad on a phone, bought on a desktop becomes two unrelated visits
without a merge. Attribution credits the last channel and zeroes out the contribution of the upper
funnel.

Personalization aimed at the wrong person. Here the shopper sees the error. Under-merging wipes
the accumulated profile: a regular customer is treated as a newcomer. Over-merging serves another
person’s interests — which reads as a site malfunction and undermines trust in recommendations
generally.

Duplicate communications. One person in two profiles receives the same email twice, and
frequency caps are counted per profile and never fire.

Implementation checklist

  1. Map the sources: where customer records are created and which identifiers each source passes —
    site, app, point of sale, loyalty programme, first-party data from
    the CRM.
  2. Normalise values before comparison: email case and whitespace, one phone format, address cleanup.
  3. Switch on deterministic rules for email, phone and internal ID — they remove the bulk of duplicates.
  4. Wire the anonymous session to the profile at sign-in and at checkout.
  5. Measure the false-merge rate on a manual sample before putting probabilistic rules into production.
  6. Write down the field merge rules, and separately the rule that the strictest consent wins.
  7. Keep the original identifiers so a merge can be rolled back.
  8. Monitor continuously: the share of profiles with no stable identifier, merges per period, manual
    splits per period.