Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

Text Similarity Checker

How much do two texts share? Four measures and every copied passage, side by side.

Text No upload Works offline Free, no sign-up

Or open a .docx, .pdf, .odt, .rtf, .html, .md or .txt file.

Nothing leaves your device.

Options

Next steps

About the Text Similarity Checker

Paste two texts — or open two Word, PDF, OpenDocument, RTF, HTML or Markdown files — and see how much they have in common, measured four ways: the passages they share word for word, word-shingle resemblance (the Jaccard index of their 3-word runs, with how much of each is found in the other), cosine similarity of the words they use, and the edit-distance ratio of their characters.

Every shared passage of six or more words is highlighted in both texts side by side; select one to find it in the other text. It is useful for checking a rewrite against its source, comparing two versions of a document, or finding the copied parts of two essays. It compares only the two texts you give it — it does not search the web — and everything runs on your device.

How to use it

  1. Paste the first text into Text A and the second into Text B, or use Open file (.docx, .pdf, .odt, .rtf, .html, .md, .txt). Nothing is uploaded.
  2. Read the headline: the share of Text A’s words that sit in passages also found in Text B, then the four measures underneath.
  3. Change Shared passages of at least to look for shorter phrases (3–5 words) or only long copied stretches (10+ words).
  4. Look at Side by side: each colour is one shared passage; select it to jump to the same passage in the other text.
  5. Copy report or Download report for a plain-text summary with every measure and a numbered list of the shared passages.

Examples

A paraphrase of a paragraph (the sample)
Input
Text A: 65 words on rainwater harvesting · Text B: a 60-word rewrite of it
Result
Shared passages: 4 (10, 6, 12 and 10 words) — 58.5% of Text A, 63.3% of Text B
Resemblance (3-word shingles): 36% · A found in B: 50.8% · B found in A: 55.2%
Cosine: 0.909 · Edit-distance ratio: 60.9% (153 edits)

The reworded sentences still share long runs such as “a small storage tank can make a real difference to a household”.

Same words, different order
Input
Text A: dog bites man · Text B: man bites dog
Result
Cosine: 1.000 (the same words) · Resemblance: 0% (no 3-word run in common)

Cosine ignores word order; shingles and shared passages do not.

Common uses

  • Teachers comparing two submitted essays, or an essay with the source it may have been copied from.
  • Writers and editors checking how far a rewrite has moved from the original, or which sentences survived.
  • SEO and content teams checking two pages for duplicate content before publishing.
  • Comparing two versions of a contract, policy or report when a line-by-line diff is too noisy.

How each measure is calculated

Words are compared after Unicode normalisation, with capitals ignored by default; punctuation and spacing never count.

  • Word-shingle resemblance — the w-shingles of a text are its runs of w consecutive words (w = 3 by default). Resemblance = |S(A) ∩ S(B)| ÷ |S(A) ∪ S(B)|, the Jaccard index of the two shingle sets, and containment of A in B = |S(A) ∩ S(B)| ÷ |S(A)|, as defined by Andrei Broder in “On the resemblance and containment of documents”.
  • Cosine similarity — each text becomes a vector of word counts; cosine = Σ aᵢbᵢ ÷ (‖a‖ ‖b‖). It is 1 when both texts use the same words in the same proportions and 0 when they share none.
  • Edit-distance ratio — 1 − d ÷ (length of the longer text), where d is the Levenshtein distance (V. I. Levenshtein, 1966): the fewest single-character insertions, deletions and substitutions that turn one text into the other. Runs of spaces and line breaks count as one space.
  • Shared passages — greedy string tiling, the method used by plagiarism detectors such as JPlag: every longest run of identical words is found, the longest are taken first, and no word belongs to two passages, so a sentence repeated in one text is matched only once.

Reading the results

There is no single “similarity score” that fits every purpose. Shared passages and “found in” show copying: a short quotation pasted into a long essay is fully found in the essay while the resemblance stays low. Resemblance shows how alike two texts of similar length are overall. Cosine shows whether they use the same vocabulary — two articles on the same topic score high even with no sentence in common. A high score is a reason to look at the highlighted passages, not proof that anything was copied: quotations, common phrases and set wording (legal clauses, recipes) are shared legitimately.

Limitations

  • It compares only the two texts you provide. It does not search the internet or a database, so it is not a web plagiarism checker.
  • Matching is word for word: synonyms, changed word endings and translated text are not recognised as the same, so heavy paraphrasing lowers every measure except cosine.
  • The edit-distance ratio is worked out for texts up to 60,000 characters each; the other measures have no such limit (up to 2 million characters per text).
  • Highlights are drawn for the first 100,000 characters of each text; the passage list and the report always cover the whole text.
  • Scanned PDFs have no text layer: make them searchable with OCR PDF first. Old .doc files must be saved as .docx.

Privacy

Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot. Opening a PDF downloads the PDF engine and an English word list from MySmartCoPilot once; your texts and files never leave your device.

Frequently asked questions

Is this a plagiarism checker?

It is a plagiarism checker between two texts you already have: it shows exactly which passages they share. It cannot tell you whether a text was copied from somewhere else on the web, because it does not search the web.

Which number should I look at?

To find copying, look at Shared passages and the highlights, and at found in for the shorter text. To see how alike two versions are overall, use Resemblance. To see whether two texts are about the same thing, use Cosine.

Why is cosine similarity high when few passages are shared?

Cosine compares which words are used and how often, ignoring their order, so two texts on the same subject — or a heavily reworded copy — score high even when no long run of words is the same.

What shingle size and passage length should I use?

3-word shingles suit paragraphs and essays; use 4–5 words for long documents, where 3-word runs repeat by chance. For passages, 6 words is a good default: shorter runs (3–4 words) are often common phrases such as “at the end of”.

Does it work for Hindi and other languages?

Yes. Words are found with the browser’s Unicode word rules, so Hindi, Tamil, Bengali, Arabic, Chinese and other languages are compared word by word in the same way as English.

Are my texts uploaded?

No. Both texts are compared in your browser and never sent anywhere. Opening a PDF downloads the PDF engine once; your file is read on your device.

Quick answers and tool search

Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.