Back to Blog

Industry Guides

Metadata in Legal Discovery: Why Lawyers Must Scrub Files Before Filing

Hidden metadata in Word docs and PDFs has cost law firms cases and sanctions.

· 9 min read

In the legal profession, a single overlooked detail can change the outcome of a case. Most attorneys spend hours reviewing every comma in a contract or every citation in a brief, yet routinely send out documents containing a hidden layer of information they have never inspected: metadata. This invisible data, embedded inside every Word document, PDF, and Excel spreadsheet, can include the names of every person who edited the file, deleted text from earlier drafts, internal comments meant only for co-counsel, and even billing notes. For law firms, learning how to properly remove metadata from legal documents is no longer a niche technical concern. It is a core duty of competence under the Model Rules of Professional Conduct.

Courts across the United States have already addressed metadata in opinions, ethics boards have issued formal opinions on the subject, and malpractice insurers have begun asking firms what scrubbing procedures they use. The risk is concrete. A pleading sent to opposing counsel with track changes intact can disclose strategy. A settlement draft circulated with hidden comments can reveal a client's true reservation price. And a Word file produced in discovery without scrubbing can hand the other side a roadmap of your internal deliberations.

What Metadata Actually Lives Inside a Legal Document

Metadata is data about data. Inside a typical Microsoft Word file, that includes the author name, the company name registered to the software license, the full edit history, the time spent in each session, deleted text preserved in track changes, comment bubbles, hidden text, and embedded objects from other files. Inside a PDF, metadata can include the original author, the application used to create the file (such as the specific document management system), creation and modification timestamps, and any annotations or form fields that were never flattened.

A surprising amount of this information is generated automatically and silently. When a paralegal opens a template, types a client's name, and saves the file, the template's original author may still appear in the properties. When a partner reviews a brief and rejects a paragraph, that paragraph can remain in the file's revision stream long after the visible text is gone. When a document is converted from Word to PDF using a simple print-to-PDF workflow, the converter often carries the author and title fields straight through.

Real Cases Where Metadata Caused Real Damage

The legal industry has already learned this lesson the hard way, repeatedly.

In one widely cited incident, a law firm representing a major corporation sent a Word version of a settlement proposal to opposing counsel. Opposing counsel opened the document properties and discovered comments between attorneys debating how low their client would actually go. The leverage in the negotiation shifted in minutes.

In another matter, a confidential memo was filed with the court as a PDF. The PDF was created from a Word file that contained tracked changes. The court's PDF viewer rendered only the accepted text, but the underlying file still contained the deletions. Opposing counsel extracted the deleted text using free tools and filed a motion citing the firm's internal strategy notes.

Ethics opinions in jurisdictions including New York, Florida, the District of Columbia, and the American Bar Association have all addressed the duty to avoid transmitting confidential metadata and, in some states, the duty of receiving counsel not to mine it. The exact rules differ, but the trend is clear: the sending attorney is expected to know what is in the file before it leaves the office.

Try MetaClean — clean this kind of file in seconds.

Strip EXIF, GPS, author, and edit-history metadata from photos, PDFs, and Office documents right in your browser.

Clean a file now · See what gets removed · Step-by-step guides · Pricing

The Specific Metadata Categories Lawyers Need to Strip

Before any document leaves your firm, several specific categories should be removed or reviewed.

  • Author and last-saved-by fields that may identify an associate, a paralegal, or a contract attorney who would not otherwise be disclosed.
  • Company name that may still reflect a prior firm if the lawyer recently moved.
  • Track changes and accepted revisions that preserve deleted language.
  • Comments and review bubbles that were intended only for internal eyes.
  • Hidden text formatted to be invisible on screen but fully present in the file.
  • Document statistics such as total edit time, which can suggest how rushed the work was.
  • Embedded files and linked objects that may carry their own independent metadata.
  • Custom properties added by document management systems such as iManage or NetDocuments, which can reveal matter numbers and internal file paths.

Why Built-In Office Tools Are Not Enough

Microsoft Word ships with a feature called Document Inspector that can remove many of these categories. It is a useful first line of defense, but it has well-documented gaps. Document Inspector does not always catch metadata embedded inside images placed in the document, does not scrub the metadata of files attached as objects, and does not handle PDFs at all. Adobe Acrobat has its own sanitization feature, but it requires a paid subscription, and many firms route their PDF creation through whichever tool is cheapest, leaving sanitization as an afterthought.

Worse, these built-in tools require every attorney and staff member to remember to run them every time. In a busy practice with overnight filing deadlines, that human step fails often.

A Practical Scrubbing Workflow for Law Firms

A defensible workflow has three parts: prevention, scrubbing, and verification.

Prevention starts with templates. Every firm template should be saved with generic author fields, no company name tied to a specific person, and no leftover content from the file it was cloned from. Practice management systems should be configured to strip identifying fields on save.

Scrubbing should be a mandatory step before any external transmission. The cleanest approach is to convert the final document to PDF, then run the PDF through a dedicated metadata removal tool that strips author, producer, creation date, modification date, and any annotations. For files that must be sent in their native Word or Excel format, such as during electronic discovery productions, run them through a scrubbing tool that handles Office Open XML internals, not just the surface document properties.

Verification is the step most firms skip. After scrubbing, open the file in a clean viewer and inspect the properties. Confirm that author, company, and revision history are gone. For PDFs, check that no hidden layers, form fields, or annotations remain.

A browser-based tool such as MetaCleanPro is well suited to the verification and scrubbing steps because it processes files locally in the browser. No document ever leaves the attorney's machine, which preserves attorney-client privilege and avoids the data residency concerns that arise when uploading client material to a third-party server. It handles PDFs, Word documents, Excel spreadsheets, and images in the same workflow.

Discovery Productions Require Special Care

Electronic discovery is a separate beast. In many cases, the parties agree or the court orders that metadata fields must be preserved in production, because metadata is itself evidence. In other cases, the protective order requires that certain metadata fields be stripped or redacted before production. Read the protective order carefully, then build a production workflow that does exactly what it requires, no more and no less. Producing too much metadata can waive privilege. Producing too little can be sanctionable.

When in doubt, log every scrubbing action so that you can later demonstrate to the court exactly what was removed and why.

Frequently Asked Questions

Does saving a Word file as a PDF remove all metadata?

No. The PDF inherits many fields from the source Word document, including author, title, and sometimes the full revision history if the conversion is not done carefully. A separate scrubbing step is still required.

Is opposing counsel allowed to look at metadata I sent them?

It depends on the jurisdiction. The ABA has taken the position that receiving counsel is generally permitted to review metadata in documents that were not produced through formal discovery, but several state ethics opinions reach the opposite conclusion. The safer practice is to assume opposing counsel will look and to scrub accordingly.

Can a court sanction my firm for metadata leakage?

Courts have ordered remedies ranging from disqualification of counsel to monetary sanctions in cases where privileged metadata was inadvertently produced. The risk is real, and malpractice carriers are paying attention.

Conclusion

Metadata is one of the few remaining areas in legal practice where a thirty-second step can prevent a career-defining mistake. Every law firm, from solo practitioners to AmLaw 100 firms, should treat metadata scrubbing as a required part of the filing and transmission workflow, on par with conflict checks and proofreading. The tools are available, the duty is established, and the consequences of skipping the step are well documented. To quickly and privately remove metadata from legal documents without uploading client files to a third-party server, use the browser-based tool at MetaCleanPro. Your clients, your partners, and your malpractice carrier will thank you.

Try MetaClean Pro free — remove metadata from your files in seconds.

Skip to main content