The Complete Guide to Duplicate Content: How to Find, Fix & Prevent It

  1. Home
  2. /
  3. Search Engine Optimisation
  4. /
  5. The Complete Guide to...

Table of Contents

The Complete Guide to Duplicate Content: How to Find, Fix & Prevent It

QUICK GUIDE – Duplicate content occurs when identical or substantially similar content appears across multiple URLs, either on your own website or across different websites. It isn’t automatically a Google penalty, but it can cause canonicalisation problems, dilute the value of your pages and make it harder for Google to understand which URL should appear in its search results.

This guide explains how to identify duplicate content, determine whether it is actually causing an SEO problem and choose the right solution.

By the end, you will know how to:

  • Find duplicate content across your website.
  • Check whether other websites are using your copy.
  • Identify copied manufacturer product descriptions.
  • Understand Google’s canonicalisation decisions.
  • Use Google Search Console to investigate duplicate URLs.
  • Use Copyscape to check content at scale.
  • Understand where SEOQuake still fits into an SEO audit.
  • Decide whether to rewrite, merge, redirect or canonicalise a page.
  • Prevent duplicate-content problems from returning.

What Is Duplicate Content?

What Is Duplicate Content In SEO?

Duplicate content is identical or substantially similar content available through more than one URL.

There are two main categories you need to understand.

Internal Duplicate Content

Internal duplicate content occurs when the duplication exists within your own website.

For example, you might accidentally have:

  • Several URLs displaying the same product.
  • HTTP and HTTPS versions of the same page.
  • WWW and non-WWW versions of a URL.
  • Tracking parameters creating additional URLs.
  • Product filters generating indexable pages.
  • Printer-friendly versions of pages.
  • Near-identical location landing pages.
  • Several service pages containing substantially the same text.

Some duplication is unavoidable and completely normal.

The important question is whether Google understands which URL you want it to index.

External Duplicate Content

External duplication occurs when substantially the same content appears on different websites.

One of the most common examples is e-commerce.

A manufacturer supplies a product description to 100 retailers.

All 100 retailers copy and paste it onto their websites.

You now have 100 product pages containing essentially the same information.

That does not necessarily mean 99 websites will receive a Google penalty.

The bigger problem is that those retailers have given Google very little reason to consider their version of the page more valuable than anybody else’s.

Does Duplicate Content Hurt SEO?

Duplicate content can hurt SEO, but not necessarily in the way older SEO advice suggested.

Google does not simply identify two matching pages and issue a “duplicate content penalty”.

Instead, Google’s systems attempt to identify duplicate and near-duplicate pages and determine which URL should be treated as the canonical version.

Problems occur when Google chooses the wrong URL or when several pages compete unnecessarily.

Possible consequences include:

  • The wrong URL appearing in Google.
  • An important page being excluded from the index.
  • Internal links pointing towards different versions of the same content.
  • Backlinks being split between several URLs.
  • Crawling resources being spent on unnecessary pages.
  • Search visibility becoming spread across competing pages.
  • Product pages offering little unique value compared with competitors.

For large websites, those problems can become significant.

Duplicate Content Is Not Automatically a Google Penalty

This is worth making absolutely clear.

Duplicate content itself is generally not a violation of Google’s spam policies.

Google routinely encounters duplicate content.

Consider:

  • Print versions of articles.
  • Mobile and desktop URLs.
  • Tracking URLs.
  • Products appearing in multiple categories.
  • Syndicated articles.
  • Sort and filter parameters.

Google’s challenge is identifying which URL should represent the content.

That process is known as canonicalisation.

Your job as a website owner or SEO is to make your preferred version as obvious as possible.

Step 1: Find All the URLs on Your Website

Before checking for duplicate content, you need a reliable list of URLs.

Years ago, I used methods involving Google search results and SEOQuake exports.

I wouldn’t build an audit around that process today.

A better starting point is one or more of the following:

  • Your XML sitemap.
  • A full website crawl.
  • Google Search Console.
  • Your WordPress database or CMS.
  • An e-commerce product export.
  • A product feed.
  • An existing website spreadsheet.

Start With Your XML Sitemap

Your XML sitemap is often the easiest place to begin.

For a WordPress website, you may find sitemaps for:

  • Pages.
  • Blog posts.
  • Products.
  • Product categories.
  • Locations.
  • Custom post types.

Your sitemap should generally contain the canonical pages you actually want search engines to crawl and index.

If the sitemap itself contains thousands of low-value or duplicated URLs, that is already something worth investigating.

Step 2: Crawl the Website for Internal Duplication

A proper crawler allows you to see how the website actually behaves rather than relying on what appears in Google.

Things worth looking for include:

  • Duplicate page titles.
  • Duplicate meta descriptions.
  • Duplicate H1 headings.
  • Near-identical body content.
  • Multiple URLs returning the same content.
  • Canonical tags pointing somewhere unexpected.
  • Redirect chains.
  • Parameter URLs.
  • Indexable filtered pages.

Don’t assume every duplicated title means you have a serious problem.

Use these reports as clues.

The objective is to find patterns.

If 400 pages share almost exactly the same body text with only the town name changed, that deserves considerably more attention than two contact pages sharing the same meta description.

Step 3: Check Google Search Console

Google Search Console should be one of the most important tools in your duplicate-content investigation.

Use URL Inspection on pages you are concerned about.

Look particularly at:

  • Whether the URL is indexed.
  • The user-declared canonical.
  • The canonical selected by Google.
  • The last crawl date.
  • Whether crawling is allowed.
  • Whether indexing is allowed.

User-Declared Canonical vs Google-Selected Canonical

This distinction is particularly important.

Your website might contain:

That is you telling Google which URL you prefer.

But Google can still decide that another URL is the better canonical.

If Google repeatedly selects a different canonical, investigate why.

Possible causes include:

  • Conflicting canonical tags.
  • Internal links favouring another URL.
  • The XML sitemap containing another URL.
  • Redirects.
  • Substantially identical pages.
  • The preferred URL being weaker or inaccessible.

Canonical tags are signals, not commands.

Step 4: Look for Copied Content on Other Websites

Internal duplication is only half of the investigation.

You should also find out whether your text appears elsewhere online.

There are several ways to do this.

Use a Quoted Google Search

Take a distinctive sentence from your page and search for it inside quotation marks.

For example:

“Our handmade oak dining tables are manufactured using sustainably sourced British timber”

If the exact wording appears across numerous domains, you may have found external duplication.

This is particularly useful for identifying manufacturer descriptions.

Use Copyscape

How To Use Copyscape To Check Article

Copyscape remains one of the best-known tools for checking whether text appears elsewhere online.

You can use it to investigate:

  • Product descriptions.
  • Service pages.
  • Articles.
  • Location pages.
  • Category descriptions.
  • Landing pages.

For an individual page, this process is straightforward.

For an e-commerce website containing thousands of products, you need a more efficient approach.

Step 5: Use Copyscape Batch Search for Large Websites

Checking 2,000 product pages individually isn’t realistic.

This is where Copyscape Batch Search becomes useful.

It allows large numbers of pages to be checked together rather than manually submitting each page.

You can even use an XML sitemap to supply your URLs.

A modern workflow might therefore look like:

XML Sitemap → Prioritise URLs → Copyscape Batch Search → Review Matches → Fix Important Pages

You do not necessarily need to check every URL.

Focus on commercially important areas first.

How Much Does Copyscape Cost?

Copyscape Premium currently charges according to the amount of text being checked.

The pricing is:

3 US cents for the first 200 words, plus 1 cent for each additional 100 words or part thereof.

That means the cost varies according to page length.

For a website containing thousands of URLs, prioritising pages before submitting them can save a considerable amount of unnecessary checking.

Step 6: Decide Which Pages Matter Most

Don’t treat every duplicate page as equally important.

I recommend prioritising URLs according to their commercial value.

Start with pages that have:

  • Strong search demand.
  • Existing rankings.
  • High revenue potential.
  • Good profit margins.
  • Backlinks.
  • Strong conversion rates.
  • Strategic importance to the business.

Suppose you have 3,000 duplicated product descriptions.

Rewriting all 3,000 randomly is not the best strategy.

Start with your 50 or 100 most commercially valuable products.

Improve those properly.

Measure the results.

Then continue.

Step 7: Work Out Why the Duplication Exists

Before fixing anything, determine what caused the duplication.

Different causes require different solutions.

Copied Manufacturer Content

This needs better original content.

Several URLs Showing the Same Page

This is probably a technical canonicalisation problem.

Old and New Versions of a Page

A redirect may be required.

Several Articles Covering the Same Subject

You may have content cannibalisation.

Product Filter URLs

You may need to review indexing, canonicals and crawl controls.

Location Pages With Nearly Identical Content

These may require substantial rewriting or consolidation.

Finding duplication is only the diagnosis.

The important part is choosing the right treatment.

Step 8: Choose the Correct Duplicate Content Fix

There isn’t one universal solution.

Here are the main options.

Rewrite the Page

Use this when the URL serves a legitimate purpose but the content isn’t sufficiently useful or distinctive.

Typical examples include:

  • Product descriptions.
  • Service pages.
  • Location pages.
  • Category content.

301 Redirect the Page

Use a permanent redirect when two pages serve the same purpose and only one needs to remain.

For example:

/old-seo-services/ → /seo-services/

A redirect consolidates users and search signals onto the preferred URL.

Use a Canonical Tag

Canonical tags can be useful when duplicate or near-duplicate URLs need to remain accessible but you want Google to treat one version as primary.

For example, product sorting or parameter URLs might canonicalise to the main category page.

Merge the Pages

Sometimes two weaker pages can be combined into one much stronger resource.

This can be particularly effective for blog posts targeting almost identical queries.

Noindex a Page

Some pages need to exist for users but do not need to appear in Google’s index.

Use this carefully.

Delete the Page

If a page serves no purpose, receives no traffic, has no links and duplicates better content elsewhere, removing it may be the cleanest solution.

Step 9: Improve Copied Product Descriptions

This is where many e-commerce businesses struggle.

What Are Long Tail Keywords?

Imagine the manufacturer gives you this:

“This lightweight aluminium mobility scooter features a folding frame and lithium battery.”

Twenty other retailers copy exactly the same description.

Simply changing it to:

“This folding mobility scooter has a lightweight aluminium construction and uses a lithium battery.”

hasn’t really improved anything.

Yes, the wording is different.

The information is identical.

Instead, expand the page with genuinely useful information.

Consider including:

  • Who the product is designed for.
  • Where it can be used.
  • How easily it can be transported.
  • Important measurements.
  • Maximum user weight.
  • Charging information.
  • Storage requirements.
  • Differences from similar models.
  • Practical advantages.
  • Frequently asked buying questions.

That’s how you turn duplicate manufacturer information into a useful landing page.

Step 10: Don’t Confuse Duplicate Content With Keyword Cannibalisation

This distinction matters.

Duplicate content means two pages contain substantially similar text.

Keyword cannibalisation occurs when several pages compete for substantially the same search intent.

You could have two completely unique 2,000-word articles and still have a cannibalisation problem.

For example:

/how-much-does-seo-cost/
/seo-pricing-guide/
/average-cost-of-seo/

If all three are trying to answer exactly the same question, Google may struggle to determine which one deserves to rank.

The solution might be consolidation rather than rewriting.

Step 11: Check Your Location Pages

Location pages deserve special attention.

Suppose you have pages for:

  • SEO Manchester.
  • SEO Liverpool.
  • SEO Leeds.
  • SEO Sheffield.
  • SEO Preston.

Changing:

“We provide SEO services throughout Manchester…”

to:

“We provide SEO services throughout Liverpool…”

doesn’t make the pages genuinely different.

Good location pages should contain useful, location-relevant information where it genuinely exists.

More importantly, the overall writing should not simply use the same paragraph structure and replace the town name.

If you are creating pages at scale, compare them against each other before publishing them.

Step 12: Use SEOQuake as a Supporting SEO Tool

How To Use SEOQuake

SEOQuake is still available and can be useful for quick SEO analysis.

However, the tool and Google’s search results have changed considerably since the original version of this guide.

I would no longer use SEOQuake as my main way of building an inventory of every page on a website.

Its value is now more as a supporting browser-based analysis tool.

What Does SEOQuake Do?

Use a crawler, sitemap or Search Console data for your main URL inventory.

Then use tools such as SEOQuake when you need quick page-level SEO information.

Step 13: Stop Relying on Old Google Search Tricks

One of the original versions of this guide involved changing Google to display 100 results per page.

That workflow is no longer appropriate.

You also shouldn’t treat:

site:example.com

as an accurate count of every indexed URL.

The site: operator remains useful for quick checks, but it isn’t a substitute for Search Console or a crawler.

Use it to investigate.

Don’t use it as your entire audit.

How Accurate Is SEOQuake?

Step 14: Forget About LSI Keywords

Another concept that still appears in older SEO tutorials is “LSI keywords”.

I wouldn’t build a modern content strategy around them.

Instead, cover the subject comprehensively.

If you are writing about loft conversions, relevant terms may naturally include:

  • Dormers.
  • Roof structure.
  • Planning permission.
  • Building Regulations.
  • Staircases.
  • Insulation.
  • Head height.
  • Structural calculations.

You don’t need a tool to tell you these are semantically related.

They appear naturally when somebody with knowledge of the subject writes a useful page.

Detailed content also gives you opportunities to appear for more specific long-tail searches.

Step 15: Check Internal Linking

Duplicate-content audits shouldn’t stop with the text.

Internal links send important signals about which pages matter.

Imagine you have:

/seo/
/seo-services/
/search-engine-optimisation/

If different parts of your website randomly link to all three, you are giving search engines conflicting signals.

Once you decide which URL is primary:

  • Update your internal links.
  • Update navigation.
  • Update breadcrumbs.
  • Update XML sitemaps.
  • Update contextual links.

Consistency matters.

Step 16: Monitor What Happens After the Fix

Once you make changes, monitor the important URLs.

I would look at:

  • Google Search Console impressions.
  • Organic clicks.
  • Average ranking positions.
  • Google-selected canonicals.
  • Indexed URL counts.
  • Organic landing-page sessions.
  • Conversions.

Don’t expect every change to produce an immediate ranking improvement.

Google needs to recrawl and reassess the pages.

For large websites, that can take time.

Duplicate Content Audit Checklist

Use this checklist when reviewing a website:

  • Download or locate the XML sitemap.
  • Crawl all indexable URLs.
  • Check for duplicate titles and headings.
  • Identify pages with highly similar body content.
  • Inspect important URLs in Search Console.
  • Compare user-declared and Google-selected canonicals.
  • Check important copy using Copyscape.
  • Search distinctive sentences manually.
  • Identify manufacturer-supplied product descriptions.
  • Review location pages for templated copy.
  • Check parameter and filter URLs.
  • Find pages targeting the same search intent.
  • Decide whether each issue needs rewriting, merging, redirecting or canonicalisation.
  • Update internal links.
  • Update the XML sitemap.
  • Monitor the results in Search Console.

Common Duplicate Content Mistakes to Avoid

There are several mistakes I repeatedly see.

Rewriting Everything

Not every duplicate page needs rewriting.

Sometimes the correct fix is a redirect or canonical tag.

Simply Rewording Existing Content

Different wording doesn’t necessarily equal better content.

Add useful information that improves the page rather than concentrating solely on making the wording different.

Ignoring Manufacturer Descriptions

Supplier descriptions can be factually useful, but don’t rely on them as your complete product content.

Creating Hundreds of Near-Identical Location Pages

Changing the town name isn’t enough.

Assuming Every Indexing Problem Is Duplicate Content

A page might fail because of technical problems, poor search intent, weak relevance or competition.

Diagnose before fixing.

Judging Content Solely by a Uniqueness Percentage

A 100% unique page can still be awful.

The objective is usefulness, not simply originality.

How Often Should You Check for Duplicate Content?

For a small service website, a duplicate-content review might only be necessary periodically or following major changes.

For larger websites, I would check more frequently.

This is especially important after:

  • Website migrations.
  • Bulk product imports.
  • Adding large numbers of new location pages.
  • Changing URL structures.
  • Installing filtering systems.
  • Importing manufacturer data.
  • Large-scale content additions.
  • Changing CMS platforms.

The larger and more automated the website becomes, the easier it is for duplication to appear without anybody noticing.

The Most Important Lesson About Duplicate Content

The goal isn’t to make every sentence on your website different from every sentence ever published online.

The goal is to make sure each important page has a clear purpose and provides enough value to deserve its own URL.

When auditing duplicate content, ask yourself:

Why does this page exist?

How is it different from the other pages on the website?

What does it provide that somebody cannot get from the manufacturer’s description?

Would combining it with another page produce something better?

Is Google indexing the URL I actually want?

Those questions are considerably more useful than obsessing over a plagiarism percentage.

Final Duplicate Content Action Plan

If you’ve discovered duplicate content on your website, work through it in this order:

  1. Find your URLs. Use your sitemap, crawler and Search Console.
  2. Identify duplication. Look internally and across other websites.
  3. Prioritise important pages. Start with traffic and revenue opportunities.
  4. Find the cause. Content, URL structure, parameters, migration or cannibalisation.
  5. Choose the correct fix. Rewrite, merge, redirect, canonicalise, noindex or remove.
  6. Improve internal signals. Update links and sitemaps.
  7. Monitor Google. Check indexing, canonicals, impressions and clicks.
  8. Prevent recurrence. Put checks in place before new pages are published.

That is the approach I would use for a modern duplicate-content audit.

The tools have changed enormously since I first started checking websites for copied product descriptions.

The fundamental principle hasn’t.

Give Google and your visitors a genuine reason to choose your page over every other version.

If you’re looking for SEO services in Leicester and need help finding duplicate content, indexing problems, keyword cannibalisation or technical SEO issues, get in touch with me at Search Focus. I work with businesses across the UK and focus on SEO changes designed to improve genuine organic visibility.

Book a Consultation



Are Orphan Pages Bad for SEO?

Are Orphan Pages Bad for SEO?

TL;DR - Orphan pages are pages on your site with no internal links pointing to them. They can hurt SEO because search engines struggle to find, index and rank them. Fix this by adding relevant internal links from high-traffic pages so content gets discovered, crawled...

How To Choose Keywords For SEO

How To Choose Keywords For SEO

TL;DR - Choose keywords that match your audience’s search intent by researching volume, competition and relevance. Use tools and competitor insight to refine long-tail and priority terms, then map them to the right pages and track performance to improve rankings and...

What Is Parasite SEO?

What Is Parasite SEO?

TL;DR - What Is Parasite SEO? TL;DR - Parasite SEO is a strategy where you publish content on high-authority external sites to rank quickly and tap their existing traffic. It can boost visibility fast, but works best when paired with strong content, relevant platforms...

What is SEO in Digital Marketing?

What is SEO in Digital Marketing?

SEO is short for search engine optimisation, a core digital marketing strategy used to improve a website’s visibility in search engine results so more people find it through organic (unpaid) search. It involves optimising content, keywords, site structure and...

How Important is SSL for SEO?

How Important is SSL for SEO?

SSL (HTTPS) secures your website by encrypting data and is a confirmed ranking factor for search engines. It builds user trust, protects sensitive information and can improve rankings and click-through rates. Without SSL, browsers warn visitors your site is unsafe,...