Check a PDF for what it reveals, then clean it

Scan for metadata, annotations and hidden-text indicators, then remove what you pick.

  • Scan
  • Remove what you pick
  • Re-scan + verify

Runs in this tab on your device — UnboundPDF has no server that receives documents. Stop is real: a stopped chain produces no file and leaves your originals untouched.

The chain scans your PDF for document information fields, an XMP packet, annotations and form widgets, embedded attachments, document actions and invisible-mode text, shows what it found, removes what you select, then re-scans the file it produced. These are indicators to review. No automated check can tell you a document is safe to share.

UnboundPDF is a free suite of 43 PDF and image tools that run entirely in your browser — merge, split, compress, edit text, OCR in 126 languages, redact, sign, convert and archive to PDF/A. Your document is read and written by the page on your own device; there is no document-upload endpoint in the core tools, no account, no watermark and no daily cap. Every result can be checked — with the Network tab, or with the Document Passport the Workspace writes for a chain of steps.

How it works, step by step

  1. Add the PDF you are about to send.
  2. Read the scan — every finding says what it is and where it lives.
  3. Choose what to remove: metadata fields, the XMP packet, annotations and widgets.
  4. Run the chain; it cleans, then re-scans the output.
  5. Review what the re-scan still sees before you share the file.

This chain reports indicators to review. It cannot certify that a document is safe to share, and it never says so — no automated check can.

What the scan looks at

Six things, all read from the file itself. The document information fields, using the reader behind Edit Metadata — Title, Author, Subject, Keywords, Creator, Producer. The XMP packet on the catalog, which repeats those and can hold editing history. Annotations and form widgets, counted page by page and reported by type, the walk Remove Annotations performs. Embedded file attachments in the name tree. Document-level JavaScript or an open action. And invisible-mode text, read from the real drawing operators.

Why “review these”, never “safe to share”

A tool that announces a document safe has made a claim about everything in it, including what it never examined: the visible content, the quality of an earlier redaction, what is inside an attachment, what a photograph shows in its background. This chain makes a smaller claim — here is what these named checks found — and leaves the judgement with you. A check that cannot run says “not checked”, never a quiet pass.

What gets removed, and what only gets reported

Removable, if you tick it: the information fields (cleared), the XMP packet (deleted), and annotations plus the form definition itself — deleting widgets alone leaves a document advertising fields that no longer exist, which readers flag. Reported only: attachments, document actions and invisible text. Deleting those silently would break legitimate documents — an OCR layer is invisible text, and removing it turns a searchable scan back into a picture.

The re-scan is the point

Cleaning and then trusting the cleaner is how leaks happen. The chain re-runs the entire scan against the bytes it produced and shows you that second list. If something survived — an attachment it will not touch, invisible text it will not guess at — it is on screen before you send the file, with the reason. The verification block underneath reports the ordinary structural checks too: the output re-opens, its page count is unchanged, it is not encrypted, and its text layer is still there.

If the sensitive part is visible on the page

Then this is the wrong tool and it will tell you so. Visible content is a job for Redact PDF, which removes the underlying text rather than covering it, and the mechanics of what really disappears are in can deleted PDF text be recovered. Rewriting a sentence rather than deleting it is editing PDF text. Freezing a filled form so its values stop being editable fields is Flatten PDF, and there is a whole outcome for that at Signature Ready.

Who needs this most

Anyone sending a document produced by someone else’s software. A contract from a firm’s template carries the firm’s Producer string; a proposal built on an old proposal carries the old client in its title. The role-specific walkthroughs go further: lawyers, HR and accountants.

The Document Passport

The passport records the chain, the checks and the SHA-256 of the cleaned file. The hash proves those exact bytes have not changed since it was written. It does not prove the document is clean, that anyone approved it, or that its contents are true — a passport from a scan full of findings is just as valid as one without. A record of what was measured, not a clearance.

Verifying that nothing was uploaded

Scanning a document for what it leaks, on a server, would be a strange way to protect it. Everything here happens in the browser tab: the PDF Privacy Lab shows how to confirm that with the Network tab open, and the comparison page states plainly where a server tool would still do better.

Frequently asked questions

Does a clean scan mean the PDF is safe to share?

No, and the tool will not say it is. A clean scan means these particular checks found nothing — no metadata values, no annotations, no invisible-text operators in the pages scanned. It says nothing about what the visible content reveals, what a redaction failed to cover, or what an embedded object contains. Treat it as a review aid.

What metadata does a PDF normally carry?

Document information fields such as Title, Author, Subject, Keywords, Creator and Producer, and often an XMP packet that repeats them. Author and Producer are the usual surprises — they frequently hold a real name and the software licence it was registered to.

What is invisible text, and why is it flagged rather than removed?

Text drawn in rendering mode 3 paints nothing but stays selectable and copyable. That is exactly how an OCR layer is stored on a scan, so removing it would break searchable scans for everyone. It is reported as an indicator so you can copy the page into a text editor and read what actually comes out.

Is this the same as redaction?

No. Redaction removes content you can see; this removes metadata around the content and the markup layered on it. If a name or number is visible on the page, use the redact tool, which deletes the underlying content rather than drawing a black box over it.