PDF to Text Extractor
Pull clean, copyable text out of a PDF — in reading order or with its layout kept.
Extracted text
About the PDF to Text Extractor
Open a PDF and get its text as plain, editable text you can copy or save as a .txt file. Reading order puts columns one after another and separates paragraphs with blank lines, which is what you want for quoting, translating, word counts or feeding the text to another program. Keep layout lines the text up with spaces the way it looks on the page, which keeps tables and forms readable.
The PDF is read by pdf.js (the PDF engine built into Firefox) inside your browser. The file is never uploaded, and password-protected PDFs work if you know the password.
How to use it
- Drop a PDF onto the box above, or tap Choose a PDF. The text of every page is extracted straight away.
- Pick Reading order or Keep layout. In reading order you can also join each paragraph into one line, undo words split by a hyphen at the end of a line, and drop repeating headers, footers and page numbers.
- To extract only some pages, type them in Pages, for example
1-3, 7or5-(page 5 to the end). - Edit the text in the box if you need to, then press Download .txt or Copy text.
Examples
The survey collected infor- mation from 2,000 well- known companies.
The survey collected information from 2,000 well-known companies.
A built-in English word list tells the two apart: "information" is a word, so the hyphen goes; "wellknown" is not but "well" and "known" are, so "well-known" keeps its hyphen.
Left column: lines 1–40 Right column: lines 1–40 Full-width footnote
All 40 lines of the left column, then the 40 lines of the right column, then the footnote
Common uses
- Copying a long passage from a report or e-book without fixing every line break by hand.
- Getting the words of a PDF into a translator, a word counter or a text-to-speech app.
- Pulling figures out of a statement or invoice with Keep layout, so the columns stay aligned.
- Checking what text a PDF really contains — for example before sending it to a search index or a screen-reader user.
How the text is put back together
A PDF does not store paragraphs or even words: it stores pieces of text and where to draw them (ISO 32000-1:2008, §9.10, “Extraction of text content”). This tool rebuilds the reading order from those positions:
- pieces on the same baseline become a line, and wide gaps split a line into separate pieces;
- lines that leave the same vertical gap free are treated as columns and read one after another, but only when they contain running text, so tables are still read row by row;
- a bigger gap between lines, a change of font size, an indented first line or a list marker starts a new paragraph;
- right-to-left scripts (Arabic, Hebrew, Urdu, Persian) are put in reading order, with numbers and Latin words inside them kept left to right.
Text is decoded through the PDF’s own character maps (ToUnicode), so accents, Indian scripts, Chinese, Arabic and symbols come out as real Unicode characters whenever the PDF provides them.
Headers, footers and page numbers
With Remove repeating headers & footers on, a line in the top or bottom 9% of a page (inside a normal one-inch margin) is dropped when the same text (ignoring numbers, so “Page 3 of 10” matches “Page 4 of 10”) appears there on at least half of the pages. Bare page numbers in those areas are dropped too, while lines set larger than the body text — chapter titles, for example — are always kept. The summary shows how many lines were removed, so you can switch the option off if something you need disappeared.
Limitations
- Scanned pages and photos of documents contain no text, only a picture of it. They are flagged so you can run them through OCR PDF first.
- Some PDFs use fonts without a character map; their text comes out as strange symbols in every extractor. When a font only leaves out some letters (common with Hindi and other complex scripts), those letters are missing and the result names the pages affected. OCR is the fix for both.
- Keep layout puts text left to right exactly as it is placed on the page, so Arabic, Hebrew and Urdu read correctly only in Reading order.
- Magazine layouts with text boxes, sidebars and captions, and very irregular column layouts, are read approximately — check the order of such pages.
- Values typed into form fields and comments are not part of the page text and are not extracted.
- The first extraction downloads the PDF engine once (about 0.5 MB compressed), so it needs an internet connection.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
Is my PDF uploaded anywhere?
No. The PDF is opened and read in your browser on your device, and the text is created there too. Nothing about the file is sent to a server.
Why is some text missing or shown as gibberish?
If a page is a scan, there is no text in it at all — use OCR PDF to recognise it. If letters come out as random symbols, the PDF was made with fonts that do not say which character each shape is; that text cannot be extracted reliably by any tool, so OCR is again the way out.
What is the difference between reading order and keep layout?
Reading order gives you flowing text: columns one after another and blank lines between paragraphs. Keep layout reproduces the page with spaces, so tables, forms and side-by-side columns stay visually aligned — best viewed in a monospaced font.
Can I extract text from a password-protected PDF?
Yes, if you know the password: the tool asks for it and uses it only on your device. PDFs that only restrict copying (an owner password) are opened without asking.
Does it keep bold, headings or tables?
No — plain text has no formatting. To keep headings, bold, lists and tables, convert the PDF with PDF to Word or extract tables with PDF to Excel.