What Hidden Information Is in Your Documents?
Metadata and Hidden Data in Word, Excel and PDF Files: Costs and Mitigation
August 25, 2026
Documents are far more data-rich than the visible content. A Word file, Excel spreadsheet, or a PDF is a container that carries identifying information that never appears on the page: who wrote it, who edited it, when, from which folder on which network drive, what was deleted along the way, and sometimes entire datasets that were never meant to leave the building.
That hidden payload travels with the file every time it's uploaded to an AI tool or provided to a third party. This article covers what's actually in there, potential costs of exposure, and how to properly scrub metadata from documents.
The Short Version
Three areas of focus:
- Metadata — descriptive information the software records automatically about the file itself: authors, dates, paths, edit counts.
- Hidden content — real content that is present but not displayed: hidden spreadsheet tabs, deleted-but-tracked text, comments, cropped image regions, filtered-out rows.
- Failed removal — content believed to be fully removed or obscured, but only partially covered up: black boxes drawn over live text, blurred numbers, "deleted" data still cached elsewhere in the file.
What's Hiding in a Word Document
Word records a running history of the document's life. Even after you delete the visible text, some of that history remains.
- Author and last-editor names. Who actually drafted the "partner-reviewed" document, or that a supposedly neutral draft came from the other side.
- Your organisation name, inherited from the licence on the machine where the file was created.
- The full folder path. Quiet but lurking risk. A file location path like
S:\Clients\Acme Corp\Matter 17 — Regulatory Inquiry\draft.docxleaks the client, the matter number, and the nature of the engagement. - Timestamps and total editing time. Useful for showing that a "carefully considered" analysis was produced in eleven minutes, or that a document existed before someone testified it did.
- Tracked changes and deleted text. Turning off the display of tracked changes does not remove them.
- Comments: including who wrote them, when, and comments marked resolved. Internal negotiations, strategy, and "candor".
- Hidden and invisible text: formatted as hidden, white text on white background, placed behind an image, or in a text box positioned off the edge of the page.
- Cropped images. When you crop a photo in Word, the cropped-away portion is usually still stored in the file at full size. Cropping is not deletion.
- Embedded objects. A small table pasted in from a spreadsheet frequently brings the entire underlying workbook with it, not just the visible cells.
- Custom fields from a document management system: client codes, matter numbers, billing references, confidentiality labels, all attached invisibly.1
- Mail-merge and link references pointing at a source list of names and addresses on a shared drive.
- Editing-session fingerprints. Word stamps each editing session with an identifier, which can be used to demonstrate that two separate documents were worked on in the same session, on the same machine.2
What's Hiding in a Spreadsheet
Spreadsheets are often built from other, larger files, and the larger file has a habit of tagging along.
- Hidden worksheets. A tab can be hidden normally, or set to a deeper "very hidden" state that doesn't even show up in the usual unhide menu.3 This is exactly how the largest police data breach in UK history happened — more on that below.
- Hidden rows and columns, collapsed groups, and filtered-out rows. A filter hides rows; it doesn't remove them.
- Cells disguised by formatting: white text on a white background, or a number format that displays nothing while the value sits underneath.
- Pivot table caches. A pivot table stores its own private copy of the source data. Delete the source sheet and the pivot still holds everything; double-clicking a total can sometimes regenerate the full underlying records.
- Formulas linked to other workbooks. The displayed value may be harmless while the formula behind it names another client's file and its exact location on your server.
- Query and connection settings: database server names, internal addresses, and the step-by-step logic used to pull and transform the data.
- Macros, which routinely contain hardcoded usernames, server paths, and commented-out code from previous versions.
- Conditional formatting rules, which quietly encode business logic: settlement authority thresholds, reserve levels, pricing floors.
- Comments and notes with author names attached, and revision history from shared workbooks.
- Sheet and workbook passwords, which are trivially removable and mainly serve to create false confidence.
What's Hiding in a PDF
A PDF with searchable text is layered, and the invisible layers still accompany the file.
- Black boxes over live text. The most common failure. Drawing a black rectangle over a name does not remove the name; select the page, copy, paste into a text editor, and it's often recoverable. This is what happened in the Manafort filings, in Apple v. Samsung, and in the 2025 release of the Epstein files.4
- Inherited Word metadata. Exporting to PDF usually carries the author name, title, and often the original file path straight across. Converting to PDF is not a cleaning step.
- A searchable text layer beneath a scanned image. If a document was scanned and text-recognised, the recognised text sits invisibly under the picture.
- Earlier versions of the file. PDFs can be saved incrementally, meaning previous states of the document are appended inside the same file and can be reconstructed.5
- Hidden layers that have been switched off, but remain in the file.
- Annotations, highlights, sticky notes and stamps, with the name and timestamp of whoever added them.
- Form field values, including in forms that look flattened to the user, and fields that were never displayed.
- Attached files and embedded images, which can carry their own camera and GPS data — a photograph exhibit may disclose exactly where and when it was taken.
- Signature details, revealing the signer's name, email address and employer.
- Pixelation and blurring, which are not reliable removal. Where the hidden content is short and predictable — a national ID number, an account number, a date — blurred output can be reversed by brute force.6
Actual Costs of Hidden Data
A hidden spreadsheet tab, a £750,000 fine, and a resignation
In August 2023, the Police Service of Northern Ireland answered a routine freedom-of-information request about staff numbers by publishing a spreadsheet online. The visible tab showed headcount. A hidden tab in the same file contained the surname, initials, rank and role of all 9,483 serving officers and staff.7 Only the visible sheet was checked before upload.
The consequences were severe:
- A £750,000 fine from the Information Commissioner's Office, the largest ever imposed on a UK public body, upheld on appeal.8
- The resignation of the Chief Constable.9
- Litigation by more than 4,000 employees, with estimated exposure of £24m–£37m.10
- Officers relocating their homes, cutting contact with family, and changing daily routines out of fear for their safety; nine in ten took up an offer of funded home security equipment.9
The regulator's central finding is the one to remember: the organisation had no procedure requiring anyone to check released files for hidden data. The file was published in its original working format rather than converted or rebuilt. A private-sector organisation making the same mistake would not have received the reduction applied to public bodies, and faced a theoretical maximum of £17.5 million.8
Redaction failures in court filings
- Paul Manafort's filings (2019). Counsel filed federal court documents whose redactions could be defeated by copying and pasting, disclosing information that had been confidential.4 It remains the standard cautionary example.
- Grand jury material. A magistrate judge ordered counsel to explain why they should not be sanctioned after a filed memorandum exposed protected testimony; the explanation was a "technical weakness" in converting from Word to PDF, and a journalist recovered the text simply by selecting and pasting the black boxes.11
- Apple v. Samsung. Confidential licensing terms and revenue figures were extractable from a filed PDF, published by the media, and prompted a sanctions motion.12
- A Florida family-court matter. A firm was sanctioned $30,000 for filing unredacted material containing a child's name, national ID number and medical information.12
- The 2025 Epstein files release. Documents were published with redactions that came apart under the same copy-and-paste test, generating an immediate round of professional commentary on what recipients of a failed redaction are obliged to do.13
Metadata alone
- An appellate court considered disqualifying counsel who made use of privileged information exposed through unstripped metadata. The court declined to disqualify but imposed other sanctions, and held plainly that a lawyer's duty of competence extends to understanding metadata and the technology handling it.14
- In another matter, producing counsel failed to strip metadata from a partly redacted privileged document; the receiving firm tried to use the recovered material as trial exhibits, but was limited by the court.15
- An audit of the US federal court filing system found roughly 1,600 cases containing unredacted Social Security numbers — the finding that drove today's mandatory redaction rules.4
The general exposure, across jurisdictions, runs to: regulatory fines, monetary sanctions and costs orders, filings struck from the record, orders to show cause, referral to disciplinary bodies, arguments that privilege has been waived, and malpractice claims.16
Why the Usual Fixes Don't Fully Work
The common approaches all share a weakness: they operate on the original file.
- Drawing boxes or highlighting adds a layer on top. The content stays underneath.
- Deleting visible text may leave it in tracked changes, in a cached copy, or in an earlier saved state inside the same file.
- Built-in document inspectors catch a useful amount, but they are opt-in, easy to forget under deadline, and they only look for the categories they know about.
- "Just export to PDF" carries metadata across rather than removing it, and can preserve a hidden text layer.
- Password-protecting a sheet can frustrate direct human access, but can be cracked and still carry accessible metadata before unlock.
How CamoText Handles This
CamoText does not overlay, mask, or edit your original file. It extracts the text, processes it locally on your device, and writes an entirely new output. The original file is never altered. Because the output is newly written rather than derived from the input container, there is no original-file metadata to inherit: no authors, no file paths, no timestamps, no comments, no hidden tabs, no pivot caches, no embedded links or media, no buried earlier versions. Those never make it into the new file.
This holds even when formatting preservation is enabled for Word, Excel, Markdown and rich text output: links, media and other metadata are wiped by default, precisely because they are such a common route to accidental exposure. For Word documents you can also choose whether tracked changes are accepted or rejected during processing, rather than leaving them to travel invisibly.
CamoText's core job is using and masking only the visible content, running fully offline, with no cloud, no API calls and no internet connection required. If you need the original terms back afterwards, a locally held key reinserts them on your own machine.
A Checklist
- No working files. Produce a purpose-built extract instead.
- Apply the copy-paste test to any PDF with redactions. Select all, copy, paste into a plain text editor. If the hidden text appears, it was never removed.
- Assume collapsed content remains in a spreadsheet, including the ones that don't appear in the unhide list.
- Document your procedure. Arm your staff with tools and simple instructions.
- Produce a fresh, clean file rather than treating the original and leaving hidden data.
Metadata and AI Tools
This risk sharpens considerably when documents go to AI. Uploading a file to a cloud model sends the whole container and metadata. A model asked to review a contract has no reason to mention that it also received the drafting history and the client's matter number, but that information has still left your control.
Pasting clean, newly written text is lighter, cheaper and faster for the model to process, and limits access to what's human-readable. See our guides on anonymising documents for AI analysis and preparing confidential documents for AI review.
Summary
Documents carry far more than they display. Tools that cover or obscure the original file often overlook many hidden vectors of information. Generating a clean file from scratch removes that burden.
Endnotes
- Document management systems such as iManage and NetDocuments write client and matter identifiers into a file's custom document properties and document variables. These are invisible in normal use and survive being emailed outside the firm.
-
Technically these are revision save identifiers, or "RSIDs", stored in the
settings.xmlcomponent of a .docx file. Word assigns a value per editing session, allowing two documents to be linked to the same session and machine; occasionally decisive in authorship and forgery disputes. -
In Excel's object model a sheet's visibility can be set to
xlSheetVeryHidden, which removes it from the standard Unhide dialog. It can only be revealed through the developer tools or by inspecting the file directly. - American Bar Association, The Judges' Journal: "Embarrassing Redaction Failures" — covering the Manafort filings, the Adobe technical guidance on improper redaction methods, and the Public.Resource.Org audit that identified roughly 1,600 federal cases containing unredacted Social Security numbers.
- Known as incremental updates. The PDF specification permits changes to be appended to the end of a file rather than rewriting it, so prior states of the document remain present in the byte stream and can be reconstructed by anyone who looks.
- Pixelation and blur are deterministic transformations. Where the underlying content is drawn from a small, structured space, an attacker can generate candidates, apply the same transformation, and match the result. Published depixelation attacks have demonstrated this repeatedly.
- Privacy World: "Data Breaches and Spreadsheets — How to Avoid Fines When Excelling" — detailed account of how the PSNI spreadsheet was prepared, reviewed, and published with the hidden tab intact.
- Information Commissioner's Office: "What price privacy? Poor PSNI procedures culminate in £750k fine" — see also the ICO's earlier notice of intent and BBC News on the fine being upheld.
- Personnel Today: "Breach of staff data sees Northern Ireland police service hit with £750k fine" — reporting the resignation of then Chief Constable Simon Byrne and the home-security payments made to affected staff.
- Infosecurity Magazine: "Widespread Security Flaws Blamed for PSNI Data Breach" — on the scale of the resulting litigation and its estimated cost.
- Nextpoint: "Don't Make These Mistakes with Redacted Legal Documents" — including the show-cause order over grand jury material exposed by black boxes applied during Word-to-PDF conversion.
- RedactLaw: "When Redaction Fails — High-Profile Legal Redaction Disasters and What They Cost" — covering the Apple v. Samsung licensing disclosure and the $30,000 Florida family-court sanction.
- The Tech Savvy Lawyer: "How To Redact PDF Documents Properly and Recover Data from Failed Redactions" — written in the immediate aftermath of the December 2025 release, and a useful summary of a recipient's obligations under the professional rules on inadvertently sent material.
- American Bar Association Litigation News: "Accidental Misuse of Privileged Metadata Results in Sanctions" — the court declined to disqualify counsel but imposed other sanctions and did not rule out disqualification in future, holding that the duty of competence extends to the technology used to handle electronic evidence.
- National Law Review: "Electronically Stored Information — Pitfalls and Ways to Avoid Mistakes" — discussing Hur v. Lloyd & Williams, in which unscrubbed metadata on a partially redacted privileged document led to an order barring the receiving firm from referring to it.
- RedactLaw: "FRCP 5.2 Compliance — Protecting Sensitive Information in Court Filings" — on the range of consequences, and the point that the filing party bears the redaction obligation regardless of who originally produced the document. See also the National Law Review on the perils of redaction.