OCR PDF (Make Scanned PDFs Searchable)
Recognise the text in scans and photos — and get a searchable PDF, plain text or hOCR.
Show the recognised text
Pages
Words to check (low confidence)
About the OCR PDF (Make Scanned PDFs Searchable)
A scanned page or a photo of a document is only a picture: you can’t search it, select a sentence or copy a figure from it. OCR (optical character recognition) reads the letters in the picture and turns them back into text.
This tool runs Tesseract 5, the open-source OCR engine, inside your browser. You get a searchable PDF — your original pages, untouched, with an invisible layer of recognised text on top, so you can search it, select it and copy from it in any PDF reader — plus the text as a .txt file and as hOCR (text with the position of every word). Nothing is uploaded: the files never leave your device.
How to use it
- Drop a scanned PDF onto the box above — or one or more images of pages (JPG, PNG, WebP, HEIC, TIFF, BMP, GIF, or a CBZ/ZIP of images). Images become one page each, in the order you chose them.
- Pick the language of the text — English, Hindi, Bengali, Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Urdu, Arabic or Chinese — and tick Also recognise English words if the document mixes in English. For a PDF you can also choose the pages (for example
1-3, 8) and the resolution: 300 DPI suits most printed documents; use 400 DPI for very small print. - Leave Straighten crooked scans on. Turn on Turn sideways or upside-down pages upright if some pages are rotated.
- Press Recognise text. The first run downloads the OCR engine and that language’s data; after that each page usually takes a few seconds on a computer (longer on a phone).
- Download the searchable PDF, the text or the hOCR file. Words recognised with low confidence are highlighted in the page previews so you know what to check.
Examples
contract.pdf — 12 page images, no text layer
contract-searchable.pdf — the same 12 pages, now searchable with Ctrl+F / ⌘F contract-ocr.txt — the text, with a line ----- Page N ----- before each page
notice-1.heic, notice-2.heic (language: Hindi)
A two-page searchable PDF using the original photos, and the Devanagari text as Unicode you can paste anywhere
form.jpg (language: Tamil, Also recognise English words ticked)
The Tamil text and the English words and figures on the same lines, both recognised — without the tick, the English words come out as Tamil letters
Common uses
- Making scanned contracts, receipts and old letters searchable before archiving them.
- Copying text out of a photo of a notice, a book page or a printed form instead of retyping it.
- Preparing scans for PDF to Word or PDF to Text, which need a text layer.
- Making documents usable with screen readers and translation tools.
What “searchable PDF” means here
Each page keeps its original appearance — for a PDF, the page itself is not changed at all; for images, a JPEG’s compressed picture data is used as it is (no recompression) and PNGs are stored losslessly. Hidden photo details such as the camera and GPS location (EXIF) are not copied into the PDF. The recognised words are added on top as invisible text: text rendering mode 3, “neither fill nor stroke” (ISO 32000-1:2008, §9.3.6). Every word is placed over its picture, so searching highlights the right spot and selecting copies the right words.
The invisible text uses a special font without visible shapes in which each character code is the Unicode character itself. That is why Hindi, Tamil, Chinese and other scripts can be copied and searched even though no font for them is embedded. Arabic and Urdu are stored from left to right, as drawn, which is how PDF readers expect right-to-left text; each reader puts it back into reading order in its own way (see the limitations below). Chinese text is stored without spaces between its characters, as it is written, so a phrase can be found as you would type it.
Getting the best results
- Resolution: 300 DPI is right for normal print. Very small print benefits from 400 DPI; 200 DPI is faster but less accurate.
- Straightening: crooked scans are straightened before recognition, and the text layer follows the original slant, so the page itself still looks exactly as it did.
- Rotation: with auto-rotate on, Tesseract’s orientation detection checks each page and turns sideways or upside-down pages upright — only when it is confident.
- Pages that already have text (from a word processor, or OCR’d before) are skipped by default, so they don’t get a second, duplicate text layer.
- Low-confidence words are marked in yellow in the previews. You can also add them to the PDF as highlight annotations, which any PDF reader can show, hide or delete.
Languages
Text can be recognised in English, Hindi, Bengali, Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Urdu, Arabic, and Chinese (Simplified and Traditional), using Tesseract’s standard trained models (the integer versions of its “best” LSTM models from the official tessdata set). Each language has its own recognition data of 1–3 MB, downloaded from this site only when you choose that language.
Many documents mix English into another language — forms, bills and notices in India often do. Tick Also recognise English words and both are read: Tesseract tries the document’s language first and English where that fits better. It downloads the English data too (about 3 MB) and can be a little slower.
Choosing the wrong language gives poor results, so pick the language the document is written in. Arabic and Urdu read right to left: the recognised text and the hOCR keep their reading order.
Limitations
- Handwriting is not recognised reliably — Tesseract is designed for printed text.
- Only the 13 languages in the language menu can be recognised; text in any other language or script comes out wrong.
- Chinese (or any text) set in vertical columns is not read correctly — only horizontal lines are.
- Urdu printed in the Nastaliq style, the usual one for Urdu, is often recognised much less accurately than Naskh-style print.
- In the searchable PDF, copying or searching Arabic and Urdu depends on the PDF reader: some readers change the order of the words on a line, or split it. The .txt and hOCR files don’t depend on the reader.
- Tables are recognised as lines of text, not as table cells. You can try PDF to Excel on the searchable PDF to rebuild them.
- Encrypted PDFs (with a password or editing restrictions) can be read but not rewritten here, so their searchable copy is rebuilt from images of the pages: text and drawings become pictures, and the copy has no password or restrictions.
- Recognition usually takes a few seconds per page on a computer and longer on a phone. Very long documents are best split into parts, especially on phones.
- The first run needs an internet connection to download the OCR engine and language data (about 5–7 MB, depending on the language, and 3 MB more with English words); your browser normally keeps them in its cache for later runs.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot. The first time you run OCR, the recognition engine and the data for the language you choose (about 5–7 MB, and 3 MB more to also recognise English) are downloaded from this site. Your files are never uploaded.
Frequently asked questions
Are my documents uploaded?
No. Recognition runs on your device, in a Web Worker inside your browser. Only the OCR engine and its language data are downloaded from this site — your files are never sent anywhere.
Will the searchable PDF look different from the original?
No. For PDFs the pages are left exactly as they were and only invisible text is added. For images, JPEG picture data is placed on the page without recompression and PNGs are stored losslessly (other formats such as HEIC or TIFF are converted once, at full quality); photo metadata such as GPS location is left out. If a page was turned upright, it is rotated, not redrawn.
How accurate is it?
Clean, printed text at 300 DPI usually gives very good results. Accuracy drops with blurry photos, low resolution, unusual fonts, coloured backgrounds and handwriting. The average confidence and the highlighted words show how sure the engine was, so you know what to check.
What is hOCR?
hOCR is an open standard for OCR results: an XHTML file with the text of every page, block, line and word plus its position (bbox) and confidence (x_wconf). Use it to build your own search index or to correct OCR output in other tools.
Can it read text in Bengali, Tamil or other languages?
Yes: Bengali, Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Urdu, Arabic and Chinese (Simplified and Traditional), as well as English and Hindi. Choose the language before you press Recognise text, and tick Also recognise English words when the document mixes in English. Other languages aren’t available yet.
Is anything stored in my browser?
The tool saves only its settings, such as the language you chose (and, like every tool, its name in your recently used list). The OCR engine and language data are ordinary files from this site, which your browser’s cache and MySmartCoPilot’s offline copies may keep so that later runs start sooner (see Cookies & storage). Nothing is put in IndexedDB, and your documents and their text are never stored.