Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

Document Similarity Checker (Offline)

Which of these documents share passages, and where? A matrix and side-by-side matches.

Education No upload Works offline Free preview, no sign-upIncluded in your pass Pro tool Pro pass: ₹179 for 30 days

Free preview.

  • Free preview: the summary counts and the most similar pair, with the first rows of the pairs table and the matrix and the first passages of each pair (up to 10).
  • Locked until you unlock it: download and copy.
  • Unlock: Pro pass, ₹179 for 30 days, a one-time payment that never renews.

Ways to unlock shows how to get the full result.

See passes (opens in a new tab)

Printing this result is locked in the free preview.

An offline comparison of your own documents. It compares only the documents you add, with each other, on this device. It does not search the internet, so it is not a plagiarism check against published sources, and shared text is a reason to look closer, not proof of copying.

Documents

    No documents yet. Add two or more to compare them with each other.

    Options

    Every passage this long that two documents share is found.
    Runs shorter than this never count as a match.
    Ignore text every document may share (the assignment question, a template)
    Any run of words this text shares with a document is left out of that document’s comparison.

    Next steps

    About the Document Similarity Checker (Offline)

    Add 2 to 50 documents — essays, assignments, reports or articles, as Word, PDF, OpenDocument, RTF, HTML, Markdown or plain-text files, or pasted text — and see which of them share passages word for word. Every pair is compared: a similarity matrix shows how much of each document is found in each other one, a ranked list puts the most similar pairs first, and the side-by-side view highlights every shared passage in both documents.

    Quotations, reference lists and text every document may contain (the assignment question, a template) can be left out, so the results show the copying that matters. The documents are read and compared on your device and never uploaded. It compares only the documents you add, with each other: it does not search the internet, so it is not a plagiarism check against published sources.

    How to use it

    1. Choose or drop the documents (up to 50), or paste each text and press Add text. Nothing is uploaded.
    2. Pick the shortest shared passage to report (8 words by default) and whether to ignore capitals, quotations and reference lists. Paste the assignment question into Text to ignore if every document repeats it.
    3. The comparison runs as soon as two documents are in the list, and again whenever you change them or the options. Read the most similar pair at the top, then the pairs table and the matrix.
    4. Select a pair (or a cell of the matrix) to see the two documents side by side, with each shared passage highlighted in both; select a passage to find it in the other document.
    5. Without a pass the free preview shows the summary counts and the most similar pair, with the first rows of the pairs table and the matrix and the first passages of each pair. With a pass, or after unlocking this result, Copy report, or download the report (.txt), every pair (.csv), the matrix (.csv) or one pair’s passages (.txt).

    Examples

    The four sample essays
    Input
    Load sample essays: four short essays on cooling a town in summer, 52 to 128 words each
    Result
    D1 ↔ D2: 41% of D1 is in D2 and 48% of D2 is in D1 (2 passages of 30 and 16 words) · D1 ↔ D3: 15% and 27% (one 17-word passage) · every other pair: 0%

    D1 and D3 also list the same reference, but the reference lists are left out. D1 has 128 words, of which 112 are compared.

    Ignoring quotations
    Input
    The same essays with Ignore quotations on
    Result
    D1 ↔ D3: 0% — their only shared passage was the assignment sentence both of them quoted
    How the shares are worked out
    Input
    D1: 112 words compared · D2: 95 words · 46 words of shared passages
    Result
    46 ÷ 112 = 41% of D1 is in D2 · 46 ÷ 95 = 48% of D2 is in D1

    A short document copied into a long one shows a high share for the short one and a low share for the long one, so both directions are given.

    Common uses

    • Teachers checking a class set of essays or lab reports for students who copied from each other.
    • Comparing several drafts or versions of a document to see what each one kept.
    • Editors and content teams checking a batch of articles for duplicated text before publishing.
    • Researchers comparing survey answers, interview transcripts or applications for text reused across them.

    How it works

    Each document is split into words (capitals ignored unless you say otherwise, punctuation never compared). Every run of k consecutive words (5 by default) is turned into a number, a hash, and the document keeps a sample of them as its fingerprints: in every window of hashes it keeps the smallest. This is winnowing, the method of Schleimer, Wilkerson and Aiken (ACM SIGMOD paper); their paper reports that the MOSS copy-detection service, used mostly for programming assignments, runs on it. With a window of t − k + 1 hashes, any passage of at least t words that two documents share is guaranteed to give both of them the same fingerprint, while runs shorter than k words can never match.

    Every fingerprint two documents share is then checked word by word and grown to the full length of the shared passage, the longest passages are taken first, and no word counts twice. The share of a document is the number of its compared words inside shared passages divided by all its compared words. Resemblance is Andrei Broder’s measure (“On the resemblance and containment of documents”): of all the different k-word runs in the two documents, the share that occur in both.

    What is left out

    • Quotations (when you tick Ignore quotations): text between paired quotation marks in one paragraph — “…”, "…", „…“, «…», 「…」 and ‘…’ (a single quote counts only where it cannot be an apostrophe).
    • Reference lists (on by default): everything from the last heading such as References, Bibliography, Works Cited, Literaturverzeichnis, Bibliografía or संदर्भ सूची, on a line of its own, to the end of the document or to an appendix after it. A heading near the start, such as a table of contents, is not taken as the start of the references.
    • Text to ignore: any run of k words that it shares with a document. Use it for the assignment question, a template or text you handed out.

    Left-out text is shown struck through in the side-by-side view, and a passage never runs across it.

    Reading the results

    A high share is a reason to read the highlighted passages, not proof that anyone copied. Students who answer the same question often repeat its words, set phrases appear in many texts, and correctly quoted and cited text is shared legitimately. Look at where the shared passages are and how long they are: one long passage of ordinary prose is far more telling than several short stock phrases. Choose a longer passage length (12–20 words) for long documents, where short shared phrases happen by chance.

    Limitations

    • It compares only the documents you add, with each other. It does not search the internet or any database, so it cannot tell whether text was copied from a website, a book or a document you did not add.
    • Matching is word for word: reworded passages, translations and synonyms are not found, and a few changed words split a passage into shorter ones.
    • Scanned PDFs have no text layer: make them searchable with OCR PDF first. Old .doc files must be saved as .docx; password-protected PDFs must be unlocked first.
    • Up to 50 documents, 2 million characters each and 10 million in all. The side-by-side view shows the first 150,000 characters of each document; the passage list and the reports cover everything.
    • A reference list is recognised only by its heading on a line of its own, and block quotations without quotation marks are compared like any other text.
    • In extremely repetitive text (the same phrase used well over a hundred times in both documents), only the first uses of that phrase are matched, so that the comparison stays fast.

    Privacy

    Your documents are read and compared in your browser and never uploaded or stored. Only your option choices are remembered in this browser.

    Frequently asked questions

    What do I get without a pass?

    Without a pass, Document Similarity Checker (Offline) shows the summary counts and the most similar pair, with the first rows of the pairs table and the matrix and the first passages of each pair (up to 10). Until you unlock it, the result can’t be downloaded or copied. A Pro or Premium pass, a one-time payment that never renews, unlocks the full result. The pricing page lists the passes and their prices.

    Is this a plagiarism checker?

    It checks a set of documents against each other: which of them share passages, and exactly where. It cannot tell you whether a text was copied from the internet or from a source you did not add, because it does not search anywhere else.

    What similarity percentage counts as plagiarism?

    No number decides it. A high share is a reason to read the highlighted passages: quotations, stock phrases and the assignment’s own words raise it without any copying, while one long passage copied word for word can matter more than a high share made of short phrases. Your school, university or publisher decides what is acceptable.

    How do I check a whole class’s essays?

    Choose all the files at once (up to 50), or drop them on the box together. Paste the assignment question into Text to ignore so that it does not count, tick Ignore quotations if quoting sources is expected, and read the pairs table from the top.

    Why are there two percentages for each pair?

    Because documents differ in length. If a 300-word answer copies 150 words from a 3,000-word essay, half of the short answer is shared, but only 5% of the long essay. The matrix reads across: each row is the share of that document found in the column’s document.

    What do the passage length and the fingerprint length mean?

    Every shared run of at least the passage length (t) is found and reported; runs shorter than the fingerprint length (k) are never matched. The defaults — 8 and 5 words — suit essays; use a longer passage length for long documents, where short phrases repeat by chance.

    Which file types can I add?

    Word (.docx), PDF with a text layer, OpenDocument text and slides, PowerPoint (.pptx), EPUB, RTF, HTML, Markdown and plain text — or paste text directly.

    Are my documents uploaded?

    No. Files are read and compared in your browser, and the comparison runs in a background worker on your device. Opening a PDF downloads the PDF engine once; the document itself stays on your device.

    Quick answers and tool search

    Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.