URLs with Similar Content: How to Fix Near-Duplicates
- December 14, 2023
- Duplicates, Content Issues

What this check flags
This flags URLs whose bodies are close but not identical — a similarity score high enough that the pages overlap heavily without being byte-for-byte twins. Unlike exact duplicates, there is no clean plumbing fix here. Every pair is a decision: are these one page wearing two coats, or two pages that haven't earned their differences yet? The answer is consolidate-or-differentiate, and on this site the owner's rule is enrich and merge, never blind-delete.
A real failing example
Two service pages scored ~90% similar because they were spun from one draft with the city swapped and almost nothing else changed:
/services/roof-repair-tampa/ body ≈ 90% match
/services/roof-repair-orlando/ body ≈ 90% match
→ same paragraphs, same FAQ, only the city name differsYou have two honest options. If the pages truly serve one intent and the local angle is cosmetic, merge them into one strong page and 301 the weaker URL:
keep: /services/roof-repair/ (national, comprehensive)
301: /services/roof-repair-orlando/ → /services/roof-repair/If both cities genuinely deserve a page, differentiate them for real — local pricing, permit rules, crew, project photos, a testimonial from that city — so the similarity score drops and each page earns its own reason to rank. What you do not do is noindex or delete a page just to make the number go down; you either make it distinct or fold its value into the survivor.
Consolidate vs differentiate: how to decide
| Signal | Lean consolidate | Lean differentiate |
|---|---|---|
| Search intent | Identical for both URLs | Genuinely different queries |
| Unique value on each page | Almost none | Real, page-specific info exists |
| Both pages ranking? | They trade places / cannibalise | Each ranks for its own terms |
| Business need | One page serves the goal | Both locations/products matter |
| Effort to make distinct | Nothing real to add | You can add substance today |
How to detect it on your own site
- Screaming Frog: enable Near Duplicates in Config → Content before crawling, then open the Content tab and filter to Near Duplicates. Set the similarity threshold (default 90%) and the tool lists pairs with their match percentage.
- Sitebulb: the Duplicate Content report has a "Near duplicate" band with an adjustable similarity slider so you can see how tight the overlap is.
- Google Search Console: in Performance, filter to a query and watch for two of your URLs alternating impressions — near-duplicates cannibalising each other show up as that flip-flop.
- Manual: paste both bodies into a text-diff tool. If the only deltas are a swapped noun and a reshuffled sentence, it is a near-duplicate.
How to fix it
- Pull the pair and read both pages before touching anything — the fix is editorial, not mechanical.
- If intent is identical and one page has nothing unique, merge the best of both into one page and 301 the loser.
- If both deserve to exist, add genuinely page-specific substance until the similarity score drops — enrich, don't strip.
- Keep the internal links and any backlinks by 301-ing merged URLs to the survivor.
- Re-crawl with the same threshold and confirm the flagged pairs no longer trip the near-duplicate filter.
Enrich, don't strip: what "differentiate" really means
When the decision lands on keep-both, the temptation is to reach for a thesaurus and reword the shared paragraphs until the similarity score dips under the threshold. That satisfies the crawler and helps nobody. A page reworded to look different but saying the same thing is still the same page; you have spent effort hiding the problem instead of solving it.
Real differentiation adds information that only that page could carry. For a set of location pages, that is local pricing, the specific team, permit or licensing rules for that area, project photos from that market, and a review from a customer there. For a set of product variants, it is the spec differences, the use cases each variant actually fits, and the questions buyers ask about that specific model. The moment a page holds something no sibling can, its similarity score falls on its own and, more importantly, it earns a reason to rank.
If you genuinely cannot add anything true and specific to a page, that is your answer: it should not be a separate page. Fold its value into the survivor and 301. The similarity flag is not asking you to obscure the overlap — it is asking you to justify why two near-identical pages both deserve to exist, and to either prove it with substance or consolidate.
Related checks
When the bodies match exactly rather than approximately, it is a plumbing fix — see exact same content and the consolidation guide. If only the tags overlap, start with duplicate title tags or same title and meta description. For the editorial call itself, when to consolidate vs create new content and the duplicate content guide go deeper.
FAQ
What counts as "similar" versus "exact" duplicate content?
Exact means byte-for-byte identical bodies. Similar means a high overlap — often 80–95% in crawler terms — where the pages share most text but differ in spots. The fixes differ: exact is a canonical/301 decision, similar is an editorial one.
Should I just noindex the weaker near-duplicate?
No — on this site the rule is enrich and merge, not prune. Either make the page genuinely distinct or fold its value into the survivor with a 301. Noindexing hides the symptom and throws away any equity the page holds.
Do near-duplicates cause cannibalisation?
They can. Two near-identical pages targeting the same query often trade impressions and split link signals, so neither ranks as well as one consolidated page would.
Are location pages automatically near-duplicates?
Only if they are spun from one template with just the city swapped. Local pages with real, page-specific detail — pricing, reviews, crew, permits — read as distinct and are fine to keep.
How similar is too similar?
There is no fixed cutoff, but if a page has no unique value beyond a swapped noun, treat it as too similar and consolidate. If you can point to real information only that page provides, it earns its place.
Claude Vincent is a technical SEO consultant focused on crawlability, rendering, and AI-search visibility. He writes the field guides and case studies at SEO ProCheck, with a bias toward the durable, unglamorous work that decides whether search engines and AI answer engines can actually read and cite a site.
About SEO ProCheck
Technical SEO consulting and GEO strategy with 20 years of enterprise experience. Case studies, resources, and tools for search and AI visibility.
Work With Me
Technical SEO audits, GEO strategy, site migrations, and international SEO. Hourly consulting for teams who need hands-on support, not just reports.







