Search Text in Multiple PDFs
Find a word or pattern across many PDFs, with the match highlighted on the page.
Files
| File | Pages | State | Remove |
|---|
Search
Matches
Pages that cannot be searched
These pages show pictures and no text — they are scans. Add a text layer with PDF OCR, then search the result.
About the Search Text in Multiple PDFs
Add a stack of PDFs — a folder of contracts, invoices, papers or manuals — and search all of them at once. Type a word or phrase, or switch to a regular expression, and every match is listed with its file, page and the words around it. Click a match to see the page with the match highlighted.
Searching ignores capital letters, accents and the way lines wrap by default (“cafe” finds “Café”, “data protection” finds “data” at the end of a line and “protection” at the start of the next), and can be made strict with Match case, Match accents and Whole words. Pages that are scanned pictures with no text are listed, so you know where PDF OCR is needed first. The results can be saved as CSV. The files are read on your device and never uploaded.
How to use it
- Choose your PDFs or drag them onto the box — as many as you like. Each file is read page by page; password-protected files ask for their password.
- Type what to find. Pick Words in order for a word or phrase, Any of the words for several alternatives, or Regular expression for a pattern. Results appear as you type.
- Tighten the search with Match case, Match accents and Whole words. Matches are grouped by file and page.
- Click a match to open the page with the match highlighted; press Download CSV to save the list.
Examples
data protection
contract-2023.pdf · page 4 … the Data Protection Act applies to … privacy-notice.pdf · page 1 … rules on data protection for staff …
Capitals, accents and line breaks do not matter, and a word hyphenated at the end of a line is joined up.
\b(?:INV|PO)-\d{4,}\bmarch.pdf · page 1 … invoice INV-20311 for … march.pdf · page 3 … order PO-5567 is …
The pattern uses JavaScript regular-expression syntax: \d is a digit, {4,} means four or more, \b is a word edge.
Common uses
- Finding which contract, policy or manual mentions a clause, a name or a figure.
- Checking a folder of invoices or statements for a reference number or amount.
- Looking up a term across the papers of a literature review.
- Looking for personal details (phone numbers, email addresses, ID formats) in files before you share them, then removing them with Redact PDF.
- Making a list of every page that mentions something, to send to a colleague as CSV.
How the search works
Each page’s text is read from the exact position of its characters, the same way the preview finds them, so a match is outlined on the page where it really is. Before searching, the text is folded: capital letters are ignored (unless you ask for Match case), accents are dropped (unless you ask for Match accents; this covers the accents of Latin, Greek and Cyrillic letters and the vowel points of Arabic and Hebrew, while the vowel signs of Devanagari, Tamil, Thai and similar scripts are part of their letters and always count), runs of spaces and line breaks become one space, and a word split by a hyphen at the end of a line is joined up (“inter-” and “national” become “international”).
Words in order looks for the phrase as typed. Any of the words finds every page with at least one of the words. Regular expression uses the JavaScript syntax; a pattern that is not valid is explained instead of searched. Whole words keeps “cat” from matching “concatenate”.
Scanned pages
A scan is a picture of a page, with no text to search. The tool reads whether each page shows pictures and no text, and lists those pages per file. Run those files through PDF OCR first — it adds an invisible text layer — and add the result here. Pages that already have an OCR text layer are searched like any other.
Patterns that take too long
Some regular expressions can take almost forever on certain text (nested repeats such as (a+)+$ are the classic example). The search runs in a separate worker that is ended when it stops responding for five seconds, and the message says so; your files stay loaded and you can change the pattern.
Limitations
- Only text that is really in the PDF can be found. Scans need OCR first; text drawn as shapes (some logos and old maps) and text in images is not text.
- Form field values, comments and other annotations are not searched: only the text printed on the page.
- The order of the words follows the order the file draws them in. In multi-column layouts and tables, a phrase that runs across columns may not be found as typed.
- At most 20,000 matches are listed (500 per page). The message says when a search stopped early; narrow the search to see the rest.
- Everything is kept in your browser’s memory. The limit is about 150 million characters of text (thousands of pages); for very large libraries search in groups.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
Are my PDFs uploaded to search them?
No. They are opened and read by your browser on your device, and the search runs there too. Nothing is sent to MySmartCoPilot or anyone else.
Why does a scanned PDF show no matches?
A scan has no text, only a picture of each page. The tool flags such pages so you know. Add a text layer with PDF OCR and search the result.
What regular-expression syntax does it use?
The one JavaScript uses (ECMAScript), with Unicode support: \d, \w, \s, [a-z], groups, alternatives with |, {n,m}, \b and lookahead/lookbehind. Case and accent settings apply to the pattern as well.
Why did the search stop with “took too long”?
The pattern made the search engine backtrack almost without end on some text, which happens with nested repeats like (a+)+. Make the pattern more specific, or use an exact phrase.
Can I search a folder?
Yes: drag the folder onto the box and every PDF inside is added. Other files in it are listed as not PDFs and skipped.