PDF Sanitizer
Find scripts, attachments and hidden data in a PDF, then save a cleaned copy.
What is inside
This tool lists what a PDF contains. It is not an antivirus: it cannot tell whether a file is harmful, and a list with nothing in it does not make a file safe. Nothing in the file is run, opened or displayed while it is read.
Choose what to remove
Tick what the cleaned copy should leave out. Your original file is not changed; the cleaned copy is saved as a new file.
The saved copy, read again
| Kind | Original | Saved copy | Result |
|---|
Report
Keyword counts, in the layout of the pdfid tool
How often the names that matter for active content appear in the file’s objects. A count above zero is not a verdict by itself; the lists above say what each one does.
| Keyword | Count | Written with #xx |
|---|
About the PDF Sanitizer
Choose a PDF and see what is inside it without opening it: JavaScript, actions that start on their own, actions that launch a program or send a form to a web address, embedded files, XFA forms, media players, comments, hidden layers, typed-in form data, metadata and page thumbnails. Every kind is listed with where it sits in the file and what it says.
Then tick what should go and save a cleaned copy. The copy is read again from scratch and compared with the original, so you can see that what you chose to remove is really gone. The tool only reads the structure of the file: nothing in it is run, opened or displayed, and the file never leaves your device.
It answers “what is in this file?”, not “is this file harmful?”. It is not an antivirus.
How to use it
- Choose the PDF or drag it onto the box. A password-protected file has to be unlocked first.
- Read What is inside. The headline says whether any active content was found, and each box below lists the items with the page or place they belong to. Open Show the … to read them.
- Pick Active content, Everything found or Nothing to tick a set of boxes, or tick the boxes yourself.
- Press Save cleaned PDF. The copy downloads, and a table shows what the original held, what the copy holds and whether each kind you chose is gone.
Examples
invoice.pdf (a form with a script, an attachment and author details)
Active content found
This PDF contains 1 script, 1 automatic action and 1 embedded file. It also holds hidden or personal data: 2 comments and 3 metadata items.
Automatic action: Runs when the document opens: JavaScript (app.alert("Welcome")) · Document
Embedded file “notes.txt” · Document · 1,204 bytes · text/plain
Document information: Author · Document · A. ExampleEach entry names the item, says where it is (a page, the document, a form field) and shows what it holds. Scripts are shown as text only; they are never run.
Kind Original Saved copy Result JavaScript 1 0 Removed Embedded files and attachments 1 0 Removed Comments and markup 2 2 Kept (not ticked)
The saved file is inspected again by the same code, from the bytes, not from a list of what was meant to be removed.
Common uses
- Looking inside a PDF from someone you do not know, without opening it in a reader, before you decide what to do with it.
- Removing scripts, attachments and comments from a file before you send it to a client or publish it.
- Taking the author name, the program that made the file and the GPS position stored inside photos out of a document.
- Finding content that sits on layers that are switched off, so it does not travel with a contract or report by accident.
- Clearing the answers from a filled-in form so the file can serve as a blank template again.
- Checking what a form does: whether its buttons send the answers to a web address.
What the tool looks at
- JavaScript in the document, on buttons and fields and in XFA forms (ISO 32000-1 §12.6.4.16).
- Actions that start on their own: the open action and the document, page and field events (§12.6.3).
- Launch and external-file actions: start a program, open another file, read data into a form, or use an address type such as
file:orjavascript:(§12.6.4). - Form submission: buttons that send the answers to a web address (§12.7.5.2).
- Embedded files: attachments, paperclip icons and portfolios, with a note on files that look like programs (§7.11.4).
- XFA, media (video, audio, 3D, Flash), comments, hidden layers (§8.11), form data (§12.7), metadata (§14.3), thumbnails (§12.3.4), and content left over from earlier saves.
The lists show the first 100 items of each kind; the counts are always complete.
How it compares with pdfid
The pdfid tool counts certain keywords in the text of a file. That is quick, but it cannot see inside compressed object streams, where many programs keep their dictionaries. This tool parses the file instead, including object streams, cross-reference streams and earlier revisions, and resolves names written with #xx escapes (/J#53 is /JS). Under Report a keyword table in the layout of pdfid shows the counts, and a separate column counts the names that were written with escapes, which is a way to hide a keyword from a simple scanner.
What removing does
Removing takes the item out of the file, not just out of sight: every object that nothing refers to any more is dropped before the copy is written, so an attachment or a script does not stay behind as an unused object.
- Scripts and actions disappear; the link or button they belonged to stays, and does nothing.
- Embedded files go with their paperclip icons.
- Comments go; links and form fields stay.
- Hidden layers: what they draw is cut out of the page content, while colours, fonts and clipping set inside them are kept so the rest of the page is unchanged. A layer is taken out of the layer list only when nothing is left on it: a block that starts or ends in the middle of a line of text, or content the tool cannot read, is left in place with its layer still switched off, and the tool says so.
- Form data is emptied, and the fields stay. An XFA form keeps a second copy of the answers in its XFA definition, so tick that too; the page reminds you.
- Metadata is emptied, including the data inside JPEG photos (Exif, GPS, XMP, comments); the photo itself is not re-encoded.
- Earlier versions and leftovers are dropped because the copy is written anew from the current version.
Why it is not an antivirus
A list of what a file contains is not a verdict. A PDF can be harmful without any of these items, for instance by damaging its own data so that a certain reader misbehaves, and a file full of these items can be harmless (many forms contain scripts that format or check what you type). If you doubt a file, do not open it, and ask your security team or scan it with an up-to-date antivirus program. This tool helps you see and reduce what a file carries.
Sources
- ISO 32000-1:2008, Document management — Portable document format — Part 1: PDF 1.7: §7.11.4 embedded file streams, §8.11 optional content, §12.3.4 thumbnail images, §12.6 actions, §12.7 interactive forms, §14.3 metadata.
- Didier Stevens, PDF tools: pdfid, the keyword triage tool whose counting this page follows.
Limitations
- It is not an antivirus. A file with nothing listed here can still be unsafe, and the tool cannot tell whether a script does something harmful: scripts are shown as text and never run.
- A password-protected or restricted PDF has to be unlocked first with Unlock PDF; the tool cannot read inside encrypted data.
- It reads the structure of the file. White text, tiny text, text covered by a shape and text inside images are ordinary page content and are not found here. To remove text for good, use PDF Redact.
- The cleaned copy is a new file. Digital signatures no longer validate, but they are not removed: a signature stays in the file, with the signer’s name and certificate. A PDF/A or PDF/X identification goes with the metadata, and the “fast web view” layout is not kept.
- Hidden-layer content that starts or ends in the middle of a line of text is left in place and reported, because cutting it out would move the text around it. The same goes for content the tool cannot read to the end or does not clean (a pattern, the picture of a note, a damaged page): the layer then stays switched off, and the table of the saved copy shows it as still found.
- A layer counts as hidden when the file switches it off in its default settings. One that is off on screen but set to print (print-only marks, say) is treated like any other hidden layer, so removing hidden layers takes it out of the printout as well.
- Scripts that are written in an unusual way are listed as scripts but not decoded or explained.
- Metadata is looked for in the places programs keep it: the document properties, XMP packets, private application data, photos, and the information of combined files and placed artwork. A name written straight into the drawing instructions of a page is page content, like any text, and is not found.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot.
Frequently asked questions
Does the tool run the scripts or open the attachments?
No. It reads the file as data: scripts are shown as text, attachments are only counted and named, and no page is rendered. Nothing in the file is run, opened or fetched, and no link is followed.
If nothing is found, is the PDF safe?
Not necessarily. The tool lists the kinds of content named above. It does not scan for malware, and a crafted file can misuse a flaw in a PDF reader without containing any of them. Keep your reader up to date and treat unexpected files with care.
What is the difference between “Active content” and “Everything found”?
Active content ticks JavaScript, actions that start on their own, launch and external-file actions, form submission, embedded files, XFA and media: things that can do something when a file is used. Everything found also ticks comments, hidden layers, form data, metadata and thumbnails: data that stays in the file unseen.
Will removing JavaScript break my form?
The fields, their values and the look of the form stay. A field that worked out its value or checked what you typed with a script stops doing so, and a button that ran a script does nothing. If the form only works through its XFA definition, see the XFA entry: that definition is removed only when you tick it.
Why is the saved file smaller, or larger, than the original?
Removed items and anything that is no longer used take space out. The file is also written again from scratch, so the way it is packed can differ a little from the original.
Is my PDF uploaded?
No. The file is read and cleaned by your browser on your device. Nothing is sent to MySmartCoPilot or anyone else.