What is duplicate content?
Duplicate content is content that exists more than once on the internet, either within one website or across several websites. That is the case, for example, when an article, a blog post or a product description is published on several web pages word for word or in a very similar form.
It is worth keeping in mind that duplicate content is not always created on purpose or for malicious reasons. It often results from technical inconsistencies or from mistakes in content management.

What types of duplicate content are there?
-
Internal duplicate content refers to content that appears more than once within a single website or domain. A typical example are product pages in online shops, where similar or identical product descriptions show up on several pages. Technical factors such as running an HTTP and an HTTPS version side by side, or having a “www” and a non-“www” version of a page, can cause internal duplicate content too. The same happens when one page is reachable through several URLs without a clear, canonical version.
-
External duplicate content occurs when identical content appears on different domains or web pages. That can happen when content is copied without permission, or when articles and blog posts are published on several platforms to reach a wider audience. Syndicating content or running guest posts on several websites can lead to external duplicate content as well.
What effect does duplicate content have on SEO?
Search engine optimization aims to place web pages in the top positions of the search results. Duplicate content can do considerable harm to that effort:
-
How search engines rate it: Search engines, Google above all, try to show users the most relevant and most unique content. Content that exists several times can therefore be rated as less valuable.
-
Competition for rankings: When several copies of one piece of content exist, they compete with each other for the rankings in the search engines. The result can be that none of the duplicates reaches a strong ranking.
-
Link equity: Link equity is the value a link passes on to a web page. When several versions of one piece of content exist, incoming links can be spread across those versions instead of concentrating their strength on a single page. That lowers your SEO efficiency.
How can I prevent duplicate content?
To prevent duplicate content, you can take several measures:
-
Canonical tags: With a canonical tag you tell search engines which version of a piece of content is the preferred or “canonical” one. That helps search engines tell the original from the duplicates. Our canonical tag generator creates canonical URLs for you automatically.
-
Avoid session IDs in URLs: Session IDs can make one and the same content reachable under several URLs. That should be avoided.
-
Careful content management: Make sure your content management system (CMS) does not produce duplicate content on its own, and check your website for duplicates regularly.
-
Robots.txt and noindex: If you have pages that carry duplicated content on purpose (print versions of web pages, for example), you can use the
robots.txtfile or set a “noindex” meta tag to keep search engines from indexing those pages.
Duplicate content is one of the standard checkpoints in an SEO audit, because near-identical URLs are usually where a crawler burns its crawl budget first. You will find more terms on the subject in the SEO glossary.