Hidden data

Remove Hidden Data and Metadata Before Sharing Documents with AI

A format-by-format guide to document properties, comments, revisions, attachments, hidden layers, image metadata, and post-removal verification.

Published August 19, 20267 min readReviewed against official sources

Why the visible page is not the whole file

A modern document is a container. In addition to the words and images you see, it may include document properties, comments, prior revisions, hidden text, worksheet tabs, speaker notes, embedded files, alternate representations, or code.

Metadata is data about the file: author, title, organisation, last editor, timestamps, camera model, geographic coordinates, or software used. Hidden content is broader and can include substantive material such as deleted-looking revisions, attachments, cropped image areas, or an entire hidden worksheet.

Changing the filename does not rewrite these structures. Exporting to another format can remove some items, preserve others, and introduce new metadata. The result must be inspected rather than assumed.

What to inspect by format

FormatExamples to inspectOfficial starting point
WordComments, tracked changes, headers, footers, hidden text, properties, custom XML, embedded objectsMicrosoft Document Inspector
ExcelHidden rows, columns and sheets, comments, names, links, queries, caches, properties, macrosDocument Inspector plus workbook review
PowerPointSpeaker notes, off-slide objects, comments, properties, embedded mediaDocument Inspector plus slide review
PDFMetadata, comments, attachments, hidden layers, cropped content, scriptsAdobe sanitisation tools
ImagesEXIF/IPTC/XMP, GPS, device data, thumbnails, visible background detailsMetadata viewer plus visual review
Text and dataComments, headers, keys, internal paths, schema names, filenamesContent and secret scan

Use native inspection tools deliberately

Microsoft recommends running Document Inspector on a copy because some removals cannot be undone. Read each category and its limitations. A hidden sheet may be part of a formula chain, and automatically removing it can damage the workbook.

Adobe separates PDF redaction from sanitisation. Redaction removes selected visible material after it is applied. Sanitisation is designed to remove hidden information such as metadata, comments, attachments, hidden layers, and scripts. Depending on the document, you may need both.

Images require both metadata inspection and visual review. Removing GPS coordinates does not remove an address visible on a sign, a customer name on a badge, or confidential information reflected in the background.

Recheck after removal

  1. Save to a new file and close the editing application.
  2. Reopen the exact output, preferably in a fresh process.
  3. Review document properties and hidden-content reports again.
  4. Search for names, domains, IDs, project terms, and other known values.
  5. Check comments, attachments, layers, notes, hidden tabs, and embedded objects.
  6. Confirm the document still opens and preserves only the information needed for the AI task.

Automated inspection has limits. Microsoft documents examples of content Document Inspector might not detect, and image OCR can miss or misread text. Use the report to guide human review, not to replace it.

A useful reporting vocabulary

Prefer precise status labels. No supported findings detected describes a scan result. Rechecked copy describes a completed workflow. Neither should be translated into safe, anonymous, or compliant without an independent basis.

A trustworthy tool should disclose its supported formats, items it can detect, items it can remove, items that require manual review, and tests used to validate output. Those boundaries help a reviewer decide what to do next.

Official sources

This guide uses primary sources available on August 19, 2026. Product policies and software features can change, so confirm current terms before handling sensitive material.