Back to Blog

Documents & PDFs

What Your PDF Reveals Without You Knowing — And How to Check

PDFs carry hidden metadata alongside their visible content: author, software, timestamps, and sometimes GPS from embedded images. Here's how to check.

· 7 min read

Quick answer: A PDF you create or edit carries hidden metadata alongside its visible content — your name, the software you used, the exact time it was saved, and, if it contains scanned or embedded photos, potentially GPS coordinates from those images. None of this shows up when you open the file normally. It shows up if someone checks the document properties, or runs the file through any basic metadata tool.

Most people assume a PDF is just the page you see. It isn't. Here's what's actually riding along with it, and what's worth doing about it.

What's stored in a typical PDF

Field What it reveals
Author Your real name or a username tied to your account
Producer / Creator The exact software and version used
Creation & modification dates Precise timestamps, including time zone offset in some cases
File path (in some exports) Local folder structure, sometimes including your OS username
Embedded image EXIF GPS coordinates and camera info, if the PDF contains scanned or inserted photos
XMP metadata A separate, often-overlooked metadata layer that basic "remove properties" tools frequently miss entirely

That last row matters more than it sounds like it should. Most simple PDF tools clear the standard Info dictionary — Author, Title, Subject — and stop there. XMP is a newer, XML-based metadata layer that can duplicate or contain more than the Info dictionary, and a lot of free tools never touch it. We cover the practical difference in deep clean vs. standard metadata removal.

Where this actually causes a problem

Job applications. A resume PDF exported from a personal laptop can carry a file path or author field tied to your current employer, disclosing information your cover letter doesn't.

Legal documents. Metadata showing a document was edited by multiple people, or by someone other than the named signatory, can matter in ways that have nothing to do with the document's visible text. This is exactly why law firms specifically strip metadata before filing.

Scanned documents with embedded photos. If a PDF includes a photo — a scanned ID, a picture inserted into a report — that image can carry its own EXIF data, GPS included, nested inside the PDF. Clearing the PDF's own Author field doesn't touch this; it requires cleaning the embedded image separately, which is exactly what a proper deep-scrub mode does.

Business documents shared externally. Internal reviewer names, department identifiers, and the specific software license your organization uses are all mundane individually, but add up to more operational detail than most companies intend to hand a competitor or an unknown recipient.

Try MetaClean — clean this kind of file in seconds.

Strip EXIF, GPS, author, and edit-history metadata from photos, PDFs, and Office documents right in your browser.

Clean a file now · See what gets removed · Step-by-step guides · Pricing

A reasonable way to think about the risk

This isn't really about any single field being dangerous on its own. An author name alone tells a recipient very little. The actual risk is aggregation: a name, plus a timestamp, plus a device fingerprint, plus a file path, adds up to a more complete picture than any one field would suggest — and it costs the recipient nothing to look, since the metadata is just sitting in the file already.

If a document is going somewhere you'd rather not be identifiable — a leak, a tip, an anonymous complaint — that aggregation is the actual threat model, not any one embarrassing detail.

Checking and cleaning a PDF

  1. Look at what's actually in the file first. Don't assume; a metadata inspector will show every field the PDF is carrying, standard and XMP both.
  2. Use deep-scrub mode for anything sensitive. A standard clean handles the visible Info dictionary fields. Deep scrub also removes embedded XMP streams and rebuilds the file's internal object structure, which is the part most free tools skip.
  3. Check embedded images separately if the PDF contains them. A scanned photo or inserted image inside a PDF can carry its own GPS data that a document-level clean alone won't reach.
  4. Verify the result. A tool that tells you "done" without showing you what was actually removed is asking you to trust it blindly — check the after-state, not just the confirmation message.

FAQ

Does "Print to PDF" remove the original document's metadata?

Not reliably. Printing to PDF creates a new file, but many export paths still carry over author information from the source document or the software's own account settings — it's not a dependable way to strip metadata.

Is XMP metadata something most people need to worry about?

For everyday sharing, usually not — the standard fields matter more day to day. It becomes relevant specifically when the document is going somewhere sensitive (legal, anonymous submission, compliance-audited), since XMP is exactly what basic "remove properties" tools tend to leave behind.

Can a PDF contain GPS data even if I never added a photo to it?

Only if it was generated from something that already had location data attached — most commonly a scanned or embedded image that itself carries EXIF GPS fields. A PDF built purely from typed text has no camera-derived location data to begin with.

Does password-protecting a PDF also remove its metadata?

No, those are unrelated. A password controls who can open the file; metadata is still embedded inside it regardless of encryption, and needs to be cleared separately.

Frequently asked questions

Does "Print to PDF" remove the original document's metadata?

Not reliably. Printing to PDF creates a new file, but many export paths still carry over author information from the source document or the software's own account settings -- it's not a dependable way to strip metadata.

Is XMP metadata something most people need to worry about?

For everyday sharing, usually not -- the standard fields matter more day to day. It becomes relevant specifically when the document is going somewhere sensitive (legal, anonymous submission, compliance-audited), since XMP is exactly what basic "remove properties" tools tend to leave behind.

Can a PDF contain GPS data even if I never added a photo to it?

Only if it was generated from something that already had location data attached -- most commonly a scanned or embedded image that itself carries EXIF GPS fields. A PDF built purely from typed text has no camera-derived location data to begin with.

Does password-protecting a PDF also remove its metadata?

No, those are unrelated. A password controls who can open the file; metadata is still embedded inside it regardless of encryption, and needs to be cleared separately.

Try MetaClean Pro free — remove metadata from your files in seconds.

Skip to main content