Your duplicate rate is a revenue number, not an admin problem
A six percent duplicate rate does not just look untidy. It splits activity history, misroutes leads and quietly understates the value of every account in your pipeline report.

Most teams treat it as housekeeping: something for whoever "owns the CRM" to tidy up when they get a spare afternoon. That's the wrong frame. A duplicate isn't two records for one person. It's one person's story, cut in half, with each half reporting into your pipeline as if it were the whole truth.
Where the six percent actually comes from
Duplicates don't arrive as one bad batch import. They accumulate, one small inconsistency at a time, from every point where a record can be created:
A prospect fills in a web form as "Bob Smith, Acme Ltd." A rep manually adds him three weeks later as "Robert Smith, Acme Limited" because they didn't think to search first. A list import from a conference brings him in a third time under his personal email instead of his work one. Your marketing automation platform creates a fourth record because it syncs on a different match key than your CRM does.
None of these are mistakes, exactly. They're the predictable result of having more than one door into the database and no consistent rule for what counts as "the same person." Every integration, every form, every rep with a "just add it" habit is another door. Multiply that by however many months or years the CRM has been in use, and six percent starts to look conservative.
What it actually breaks
The obvious cost is wasted admin time merging records. That's the smallest part of it.
Activity history splits. Every call, email and meeting logs against whichever record was open at the time. Look at "Bob Smith" and you see two calls. Look at "Robert Smith" and you see one email and a demo. Neither record shows the true picture, so nobody, not the rep and not their manager, can answer "where does this deal actually stand?" without manually stitching two records back together.
Leads get misrouted. Lead-routing rules typically fire on the record being created, not on the person behind it. A returning lead who books a demo on their second visit can get assigned to a different rep than the one already working the deal, because the system sees a "new" contact. Two reps end up quietly working the same company, unaware of each other, sometimes with two different quotes in flight.
Pipeline value gets understated. If a contact has two records, their deal history is probably split across both. A report that sums deal value "by contact" silently shows two smaller numbers instead of one accurate one. And if your reporting rolls up by account, a company's total spend can look meaningfully lower than it actually is, right when someone's deciding whether that account is worth investing more sales time in.
Forecasting drifts. Sales forecasts are built by summing what's in the pipeline. Duplicate deal records, the same opportunity logged twice because it started under two different contacts, inflate the number. Duplicate contact records with the value split between them deflate it. Neither error is visible in the total; it just makes the forecast wrong in a direction nobody can explain.
Marketing spends twice, and it shows. Two records for one person often means two nurture sequences running in parallel, or the same person receiving the same campaign email twice on the same day. It's a small thing individually. At scale, across a database with a six percent duplicate rate, it's a steady leak in marketing spend, and a steady, visible signal to prospects that you don't have your act together.
The maths on six percent
Take a CRM with 40,000 contact records and a six percent duplicate rate. That's roughly 2,400 records that shouldn't exist as separate people. Call it 1,200 real contacts represented twice each.
If even a fifth of those 1,200 people are currently active somewhere in a deal, that's well over 200 opportunities where activity history, deal value or lead ownership could be split, misassigned, or double-counted. You don't need a large fraction of your database to be duplicated for the effect on forecasting, routing and reporting to be material. You just need duplicates concentrated in the records that are actually moving through the pipeline right now, which they usually are, because active records are the ones being touched, updated and re-created most often.
Why it gets worse on its own
Left alone, a duplicate rate doesn't stay flat. It compounds. Every fresh interaction with an already-duplicated contact is another chance to create a third record instead of finding the other two. Every list import repeats whatever inconsistency caused the first round of duplicates, because the underlying cause (no single rule for what counts as a match, no enforced search-before-create step) is still there. A CRM that isn't actively protected against duplication doesn't hold steady at six percent. It drifts toward eight, then twelve, then a point where nobody trusts the reports enough to use them, and the real pipeline moves into a spreadsheet instead.
Fixing it: three layers, not one clean-up
1. Clean what's already there. A proper dedupe pass, using your CRM's native matching tools or a dedicated deduplication tool with clear rules for which record survives and which fields get merged rather than overwritten, needs to happen before anything else. Do this without a plan for what comes next and you'll be back here in a year.
2. Stop new duplicates at the door. Form validation that checks for an existing match before creating a new record. A single, enforced entry point for new contacts rather than five loosely-connected ones. A "search before you create" step built into the rep workflow, not just written in a wiki nobody reads. Consistent formatting rules for company names and required fields, applied at the point of entry, not cleaned up after the fact.
3. Monitor it as a number, not a project. Duplicate rate should be a metric someone actually looks at, monthly rather than annually, the same way you'd track any other data quality or pipeline health figure. A one-off clean-up fixes today's problem. A tracked number catches next quarter's before it costs you a forecast.
What a healthy number looks like
As a rough benchmark: under roughly 2% is what a well-maintained, actively governed CRM usually holds at. Between 5% and 10% is common for a CRM that's never had a structured dedup process, worth fixing, not yet an emergency. Above 10%, or trending upward, is the point where it's actively distorting your reporting and your reps' day-to-day, not just cluttering the database.
If you don't currently know which of those bands you're in, that's the actual first problem: not the duplicate rate itself, but not knowing it. It's the first number we check on every data audit, because it's the fastest way to tell whether the rest of your reporting can be trusted at all.