The technical mistakes that out a confidential source are often very small. A document property field. A printer's tracking dots. A revision history left in a Word file. A photograph's GPS coordinate. Metadata leaks have repeatedly identified people whose anonymity was the entire point of their communication — and in nearly every case, the leak did not require a sophisticated attack. It required someone to look at fields the sender forgot were there.
Three well-documented cases show how it happens, and a fourth shows what is now expected of any organisation that publishes a document on a source's behalf.
The 2003 Iraq dossier
In February 2003, the UK government published an intelligence dossier on Iraq's security infrastructure. Within days, Cambridge University researcher Glen Rangwala examined the Word document's properties and found the revision log named four civil servants — including one whose section had been lifted, lightly edited, from a graduate student's thesis on al-Qaida. The discovery did not require special tools: Word itself, set to display revision authors, listed every name. The dossier became known as the "dodgy dossier" within a week of publication, in significant part because its document properties told the story of who had touched it.
Reality Winner and the printer dots
In June 2017, The Intercept published a leaked NSA document about Russian election interference. The PDF was a scan of a printed page. Embedded in the scan, invisible to the human eye, were the Machine Identification Code dots that most colour laser printers add to every output: a grid of yellow microdots encoding the printer's serial number and the timestamp of the print job.
NSA contractor Reality Winner was arrested before the article was published. Court filings later confirmed the agency had reconstructed the printer used to produce the leak from the dots, traced it to a specific machine, identified six people who had printed from that machine in the relevant window, and narrowed it to Winner from there. The scan that the news outlet had published carried the source's identification on every page.
How the printer dots actually work
Most colour laser printers from the major manufacturers add a grid pattern of pale yellow dots to every page. The pattern encodes the printer's serial number, the date and the time of the print, in a binary encoding researchers (notably the Electronic Frontier Foundation) reverse-engineered in the mid-2000s. The dots are barely visible under normal light and obvious under a UV torch or a high-quality scan run through a contrast filter.
Stripping EXIF or document metadata does nothing to printer dots. They are physically on the paper. The only defences are to print in monochrome (most monochrome lasers do not add dots), to use a printer that does not encode them, or to not include scans of printed pages at all when source protection matters.
Cablegate and the redaction failures
When WikiLeaks began publishing US diplomatic cables in 2010, the process of redacting sensitive names was deliberately laborious because the team had learned from earlier failures. Even so, a small number of cables were released through partner outlets with PDF redactions that turned out to be drawn over selectable text. Within hours, researchers had extracted the names that had been "redacted" with copy-and-paste in a different reader.
The technical lesson was the same one the US Department of Justice had already received, twice, over the previous decade: drawing a black box on a PDF is not redaction. The procedural lesson, internalised by every publisher of leaked documents since, is that source protection has to assume incompetence on the part of the redactor and check the result with text extraction before publication.
What publishers do now
Modern source-protection workflows for journalists, NGOs and law firms typically include:
- Re-creating documents from text rather than passing on the original file. The source's Word document is read, its content typed into a new file, the original destroyed.
- Re-rendering scans through a process that obliterates printer dots: print to PDF, run through a tool that flattens to a low-DPI greyscale image, OCR the result, throw away the scan.
- Stripping all EXIF, IPTC and XMP from photographs supplied as evidence, including the serial numbers most cameras write.
- Running the final document through a metadata cleaner to remove any properties the editing software left behind.
- Verifying the result with text extraction tools before publication, treating any returned hit on a redacted name as a publication-blocking failure.
The cleaner on the homepage exists to handle the last step in that chain — the residual document and image metadata — without uploading the file anywhere. It is not a replacement for a thoughtful workflow; nothing technical is. But it removes the layer of risk that has, over and over again, put real names on documents that should not have carried them.
The cases above are public because they ended badly. The cases that did not end badly are quiet, and they almost always include someone, late in the workflow, checking what the file knew about its source.