Skip to content

What is duplicate content and how does it affect SEO?

Author: Matteo Pellegrini

Most duplicate content is accidental rather than plagiarised, so Google does not penalise it, and that is one of the SEO myths that keep coming back. Even so, duplicate content has a negative effect on your search engine optimisation (SEO) work.

Read on to find out what duplicate content is, how it affects you and how to prevent it.

What is duplicate content in SEO?

Duplicate content is identical or near-identical content that appears on different URLs. If a page contains exactly the same text as another page, it counts as duplicate content.

Duplicate content can sit within the same website or on pages of different websites.

Does Google penalise duplicate content?

"Duplicate content on a site is not grounds for action on that site unless it appears that the intent of the duplicate content is to be deceptive and manipulate search engine results." (Source)

Google

Google does not penalise duplicate content, at least when it is not deliberate.

However, deliberately copying content from another website and republishing it as your own goes against the spam policies in Google Search Essentials. Scraping can push a site lower in the search engine results pages (SERPs) or stop it from appearing at all.

How does duplicate content happen?

Most published duplicate content is unintentional. Some site owners do not even know they have created duplicate content on their own site.

Here are four ways duplicate content can appear on a site:

1. URL variations

Your site may create new URLs without you noticing when it uses session IDs or click tracking, so what should be a single URL ends up as several. This is the most frequent of the common URL problems.

A printer-friendly version of a page can also create duplicate content when other versions of a URL get indexed.

2. Site versions

Does your website have both HTTP and HTTPS versions? If so, you have created duplicates of your site or your pages. A website that is reachable with and without "www" at the start may also have generated copies of its pages or of the whole site.

3. Copied content

Content scraping means copying content from one page to another. Sometimes there is no ulterior motive. For example, two different distributors of the same brand may have product pages with very similar content.

4. Accidental duplication

Different websites can create and publish similar content. News sites, for instance, cover the same events. Different distributors of the same brand and products can have almost identical category pages.

Why is duplicate content a problem for SEO?

There are several reasons why duplicate content causes SEO problems, including:

Crawling and indexing

Crawling and indexing the web is expensive, which is one of the reasons Google has raised its content quality standards.

If you produce duplicate content at scale, you will probably see a difference (and not a good one) in how Google crawls and indexes your site, and that will hit your organic pipeline.

Rankings

In many cases Google indexes duplicate content but struggles to show it in search results. Which page should rank? Most of the time the answer is neither. Google will rank the pages poorly, and the result is little or no organic traffic.

On the other hand, when you fix duplicate content you can see a clear effect on organic results.

User experience

Duplicate content also affects user experience. If users can easily reach these duplicate pages, they start to feel lost and lose trust in your website, which drags down your wider marketing effort.

Online reputation

Sites that create duplicate content by copying other people's work end up with a poor online reputation. Reputation matters: it inspired Google's PageRank system, which works much like citations in academic research, and it is why unique content pays off.

How to find duplicate content

You can find duplicate content on your site in several ways, including:

  1. Crawl your site with a free SEO tool such as Screaming Frog and go through the results.
  2. Check the page indexing report in Google Search Console.
  3. Analyse your site with an AI SEO tool, such as ChatGPT or Claude.

How to prevent duplicate content

Now that you know what duplicate content is and how it affects SEO, here are the best practices to prevent it on your site:

  • Use 301 redirects
  • Guide search engines with correctly set canonical tags
  • Use a noindex meta robots tag
  • Avoid publishing duplicate content where you can.

Let's look at each one in turn:

Use 301 redirects

301 redirects are a good way to deal with duplicate content. When you move from HTTP to HTTPS, 301 redirects tell search engines to send anyone who requests the HTTP page to the HTTPS one.

That way, every user who wants to visit your page lands on the HTTPS version, even when they try to open the HTTP page.

Redirects are also useful when you need to merge two or more pages and point them to a single one, as long as you pick a permanent code: a 302 redirect would leave the old address in the index.

Say you have published a blog post on a topic you had already covered. You can merge the content into a single page, ideally the one that ranks higher, and then 301 redirect the other URL to it.

Guide search engines with canonical tags

Do you have a printable PDF version of one of your HTML pages?

You can tell Google that the PDF is a duplicate and that it should treat the HTML version as the original. For a PDF, which has no HTML head, you do this with a rel="canonical" HTTP header sent with the file.

Use the noindex meta robots tag

The noindex meta robots tag is a line of code you add to the <head> section of a page to tell search engines to keep it out of the index and the SERPs. The code looks like this:

<meta name="robots" content="noindex">

With this tag you keep duplicate content out of the SERPs and send traffic to the versions of the page you are actually optimising.

Avoid duplicate content where you can

If you notice that a page generates several URLs for different sessions, consolidate them into one.

And if you run a blog that you update regularly? Check your site every so often for posts on similar topics that you can merge into a single article.

Cut duplicate content and your SEO work goes further

To improve your rankings in the SERPs, your site needs a smooth user experience and useful content. Duplicate content on your site can hold your rankings back and confuse visitors.

Want to improve your SEO strategy?

The Visilay team can help you find and fix duplicate content issues while optimising your site to perform better. Get in touch to find out how we can support your SEO strategy.

Matteo Pellegrini

Matteo Pellegrini

I’m a Business Developer, and at Visilay I focus on developing data-driven SEO, Google Ads, and CRO strategies. I love historical museums, have been practicing Karate for as long as I can remember, and on weekends I enjoy exploring Italian villages in search of authentic local food.