Sanitize a PDF — see what it carries besides its pages, then strip what you choose

Updated 28 August 2026

Sanitize PDF reports what a document carries alongside its visible pages — metadata fields with values, annotations by type, embedded attachments, scripts and open actions, form-field values, and text still hidden under a cover — lets you tick what to strip, and re-checks the saved output independently, naming anything still present. Sanitising is not redaction: nothing printed on the page is changed, and hidden text is flagged with a link to Redact PDF rather than removed.

Check a PDF before sending: a findings report of metadata, comments, attachments, scripts and hidden text — strip what you tick, verified on the output.

UnboundPDF is a free suite of 48 PDF and image tools that run entirely in your browser — merge, split, compress, edit text, OCR in 126 languages, redact, sign, convert and archive to PDF/A. Your document is read and written by the page on your own device; there is no document-upload endpoint in the core tools, no account, no watermark and no daily cap. Every result can be checked — with the Network tab, or with the Document Passport the Workspace writes for a chain of steps.

  1. Open Sanitize PDF and add the document — the findings report comes first, class by class with counts.
  2. Tick what to strip; metadata, scripts and attachments are on by default, comments and form values wait for you.
  3. Download and read the check made on the output bytes — what is no longer present, and anything that survived, named per page.

Open Sanitize PDF and add the file. The findings report comes first — you decide what goes; the tool then shows, on the output bytes, what went.

What a PDF carries that you cannot see

Author fields holding a real name, the review thread nobody deleted, an attachment embedded years ago, document JavaScript, form fields still holding values — and sometimes text that is present but painted invisibly. All of it travels with every copy.

The report is the product

Six classes, counted, with pages. Then your call, with conservative defaults. Hidden text is deliberately different — flagged with no checkbox, because stripping it blindly would break every OCR'd scan; the row routes to Redact PDF and the Redaction Verifier instead. Bookmark titles are outside every class and never touched — stated on the page.

The check you can quote

After saving, independent code re-reads the output and reports per class: checked and no longer present — or still present, named per page. The vocabulary is deliberate: never "clean", never "safe", because no automated check can promise that. What it verified is the only promise worth forwarding.

Frequently asked questions

How is this different from redaction?

Redaction removes visible content; sanitising removes what the file carries alongside the pages. Hidden text is flagged here — with no checkbox — and routed to Redact PDF.

What does the final check verify?

The saved bytes are re-scanned by code independent of the code that edited them. The result says which classes are no longer present — and names anything still there. Never 'clean', never 'guaranteed'.

Why are comments off by default?

A comment thread and a filled form are usually the document's content; a tool that silently deletes them destroys work. You tick them deliberately.