The names, timestamps, device IDs, GPS coordinates and authorship trails embedded in a file are not just technical detritus. Under EU and UK data protection law, much of it is personal data — subject to the same lawful-basis, retention and disclosure requirements as any column in a customer database. Metadata under GDPR is one of the categories regulators have begun to scrutinise more directly, and the line between "harmless file properties" and "an unauthorised disclosure of personal data" is narrower than most organisations assume.
Why hidden file metadata is personal data
The UK GDPR (and the EU original) defines personal data as any information relating to an identified or identifiable natural person. A name in a Word document's creator field is personal data on its face. A GPS coordinate in a photo is personal data when combined with the file's other identifiers (filename, account that uploaded it, timestamp). A camera serial number, an email address in a comment, a UPN in a threaded comment, a Windows account name in lastModifiedBy — each is, on its own or in combination, capable of identifying a person.
The Information Commissioner's Office and several European data protection authorities have been explicit that file metadata falls within scope. The European Data Protection Supervisor has published guidance on metadata in published documents specifically. National regulators including the Spanish AEPD have issued enforcement decisions involving unintentional disclosure of personal data via uncleaned document properties. None of this is novel law; it is the application of existing definitions to artefacts most data protection programmes never thought about.
Where the practical risk shows up
Subject access requests
A subject access request entitles a data subject to a copy of all personal data the controller holds about them. If a candidate's CV ends up forwarded between recruiters, the metadata of every forwarded copy — the names of the recruiters who edited it, the agencies whose templates it passed through, the comments left in the comment pane — is, technically, the candidate's personal data. A complete response to a subject access request that returns only the visible content is not a complete response.
Public document publication
Government bodies, NHS trusts, councils and regulated entities publish documents under transparency obligations: meeting minutes, policy drafts, FOI responses, consultation papers. Publishing those documents with their metadata intact has, repeatedly, breached the data protection rights of the staff named in them — junior officials whose names appear in the creator field, contractors whose comments survive in the comment pane, third parties whose redacted names are recoverable from the underlying text stream.
The ICO's general expectation, set out in published guidance, is that public-sector publishers strip personal metadata before release as a matter of standard process.
Data sharing under contract
A controller-to-processor or controller-to-controller transfer typically specifies what personal data is shared. A spreadsheet sent to a processor for a statistical analysis is shared on the basis of its anonymised columns. The same spreadsheet's creator, lastModifiedBy, comment author names and pivot-cache rows can carry personal data the contract did not authorise. The transfer of unintended personal data via metadata is not a contractual nicety; it is a processing operation without a lawful basis.
Breach notification
Under Article 33 of the UK and EU GDPR, a personal data breach that is likely to result in a risk to the rights and freedoms of natural persons must be notified to the supervisory authority within 72 hours. A document published with personal data in its metadata, exposed to the public web for any non-trivial period, is, on a strict reading, a notifiable breach. Whether it actually rises to a notification depends on the data, the risk and the documented assessment — but the obligation to perform that assessment is real.
What a defensible metadata-cleaning policy looks like
A small set of measures, applied consistently, brings most organisations within scope of accepted practice:
- A documented policy that identifies file metadata as personal data and assigns responsibility for cleaning it before external disclosure.
- Standard cleaning steps in document and image production workflows, integrated into the tooling people already use rather than added as a separate task.
- Defined exceptions: where metadata is required to be retained (e.g. evidential continuity in legal disclosure), the basis is documented.
- Training that makes clear what metadata exists in common file formats — the Word author field, the PowerPoint embedded image, the EXIF GPS tag — rather than treating the issue as abstract.
- A technical control that strips metadata at the point of egress: an export gateway, a browser-based cleaner, or a server-side process that runs on outbound files. The control should be applied without depending on individual users remembering to use it.
How a cleaner fits into the picture
Cleaning is one control among several, but it is the one that catches the residue everything else missed. The cleaner on the homepage processes files in the browser without uploading them, which itself simplifies the lawful basis for using it: no transfer of personal data takes place during cleaning. For organisations with bulk-cleaning needs, the API-based versions perform the same operation programmatically, integrated into outbound document and image flows.
Whatever tooling you choose, the working assumption to internalise is the simple one: the metadata in a file is personal data when it identifies a person, and most files contain more of it than the people sending them know. A metadata-cleaning step in your release process is not paranoia. It is the routine application of a rule the regulators already wrote a decade ago.