Duplicate Text Checker
Compare two drafts side by side to measure textual overlap, inspect repeated wording, and decide whether the pages need consolidation, clearer differentiation, or no change at all.
What this tool checks
Overall Similarity Score
Combines word-level and phrase-level overlap into one comparison score. Under 50% indicates lower measured overlap, 50-80% indicates substantial similarity that deserves review, and above 80% indicates very close textual resemblance. These are tool review bands, not Google penalty thresholds, so use them to prioritize inspection rather than to predict indexing or ranking outcomes.
Word-Level Overlap
Compares the unique words shared by both texts using Jaccard similarity. This helps distinguish broad topical similarity from more direct textual reuse, which is useful when deciding whether two drafts genuinely serve different purposes or mostly cover the same material.
Phrase-Level Overlap
Compares shared 3-word phrases between the texts to surface passages that use closely matching wording and sequence. Phrase overlap can reveal copied or lightly reworked sections that a simple unique-word comparison may not make obvious.
Why Duplicate Text Deserves Review
Substantial overlap between pages can make a site's content architecture harder to interpret and can leave multiple URLs serving nearly the same reader need. The practical decision is not to rewrite every similar passage, but to determine whether each page has a distinct purpose, useful original information, and a clear reason to exist. Where pages are genuinely redundant, consolidation, canonicalization, or a more differentiated content strategy may be appropriate depending on the publishing situation.
Common issues this tool detects
Multiple pages can target almost the same intent while repeating most of the same explanations. Review whether each URL answers a meaningfully different question; if not, consider consolidating the useful material into a stronger preferred page.
Product variants may share 90%+ of their descriptive copy when only a small attribute changes. That similarity is not automatically a problem, but pages should still provide enough variant-specific information to justify separate URLs when separate pages are necessary.
Syndicated or republished material can appear across multiple domains. When duplication is intentional, use an appropriate canonical or publishing arrangement when supported by the use case, and verify that users can still identify the original or preferred version.
Repeated navigation, legal text, company descriptions, or templated location copy can make many pages look similar. Treat 20% as a tool review reference rather than a search-engine limit, and focus on whether the main page content provides enough distinct information for its actual purpose.
Frequently asked questions
How is the similarity score calculated?
The comparison uses Jaccard similarity at the unique-word level and also checks shared 3-word phrases. The combined view helps separate broad topical overlap from closer wording reuse so you can see whether two texts merely discuss the same subject or repeat substantial passages.
What should I do if two pages have high similarity?
If similarity exceeds 50%, review whether both pages serve distinct intents. Options include (1) consolidating genuinely redundant material into a preferred page and using a 301 redirect where a permanent URL move is appropriate, (2) rewriting one page around a clearly different purpose with original evidence or examples, or (3) using a canonical when separate but substantially similar versions legitimately need to remain accessible.
Does Google penalize duplicate content?
Duplicate content is not automatically a manual-action issue. Search systems may choose among similar versions or cluster duplicates, so the practical concern is whether the preferred page is clear and whether each indexed URL offers distinct value. Deliberately deceptive or scraped content raises different quality and policy considerations.
What percentage of similarity is too much?
Treat the thresholds as comparison bands rather than universal SEO rules. Under 50% indicates lower measured overlap, 50-80% calls for closer review, and above 80% indicates very similar text. The right action still depends on search intent, page purpose, necessary boilerplate, and whether the shared wording is substantively useful.
What is keyword cannibalization?
Keyword cannibalization is a practical label for situations where multiple pages on the same site target substantially the same query intent and compete for attention. Similar wording can be one clue, but the more important question is whether the URLs have overlapping purposes and whether one page would serve users more clearly.
How is this different from a plagiarism checker?
This tool compares the two texts you provide and measures similarity between them. A plagiarism checker generally attempts to find matching material across a broader corpus. Use this comparison for draft review, internal content audits, migration work, or deciding whether two pages are sufficiently distinct.
Should I use canonical tags for similar pages?
Use a canonical when substantially similar pages legitimately need to remain available but one URL should be presented as the preferred version for indexing. Do not use canonicalization as a substitute for fixing pages that should actually be merged, redirected, or made meaningfully distinct.
How much unique content should each page have?
A previously published rule of thumb suggested aiming for 60-70% unique content and reviewing pages where 30-40% is repeated boilerplate. Those figures are not universal Google thresholds. Judge uniqueness by whether the page contributes distinct, useful information for its intended query and audience rather than by percentage alone.