The black box you draw across a name in a PDF is, for most software, just paint. The text underneath it is still there, sitting in the same coordinates it always did. Anyone who copies the page into a different reader, or extracts the text stream directly, gets back the words you thought you had hidden. PDF redaction failures keep showing up in court filings, government releases and corporate disclosures because the tooling and the workflow disagree about what redaction means.
The mistake is rarely technical incompetence. It is a misreading of what a PDF actually is.
How a PDF stores text
A PDF page is a stack of independent layers. There is a content stream of drawing commands, an embedded font subset, an internal text positioning grid, and any number of optional annotation layers on top. When you open a document and see "Mr Andrew Smith, 12 Maple Lane", the reader has executed a command like (Mr Andrew Smith, 12 Maple Lane) Tj at a specific position.
If you draw a black rectangle on the page using an annotation or a comment tool, you have not changed the underlying text command. You have added a new shape on a higher layer. The original Tj command is still in the file, the search index still hits it, and copy and paste still pulls it out.
Where the leaks appear
Selectable text under the box
Most failures are this simple. Open the redacted PDF in any reader, drag your cursor over the black box, and the highlighted text shows up in your clipboard exactly as the author typed it. Court filings, agency releases and corporate disclosures have leaked this way for two decades. The pattern repeats because each new generation of staff inherits the same misconception about what a black rectangle does.
Image-based redaction over searchable text
A more subtle failure: the visible page is a rendered image, but the PDF still carries an OCR text layer underneath it for accessibility. The image hides the words from your eyes; the OCR layer hands them over to anyone with pdftotext.
Bookmarks and outlines
The body of a 200-page report can be carefully redacted while the table of contents at the top, generated automatically from heading styles, still lists "Section 4.3: Allegations against [the name you redacted]". Bookmarks pull from the same source.
Document properties
The /Info dictionary and the XMP packet carry the original Title, Author, Subject and Keywords as they were typed. A document titled "Smith dismissal — final" with all visible references to Smith blacked out still says Smith on every reader's title bar.
Embedded objects and attachments
PDFs can contain attached files (the paperclip icon in some readers). A press release exported "for safety" sometimes ships with the original Word source still bolted on inside.
Form fields and JavaScript
Interactive PDFs carry their form data and any embedded JavaScript in clear. A redacted form that retained the original field values, or a document that ran client-side scripts during editing, leaves both behind unless the file is explicitly flattened. Flatten before publishing and the page becomes a static rendering of what it shows; skip the step and the underlying field tree is part of the file.
What real redaction has to do
True redaction modifies the underlying content stream. It deletes the original Tj command, removes the matching range from the text layer, strips the corresponding bookmarks, rewrites the OCR layer if there is one, and clears anything in the document properties that referenced the removed material. That is a destructive edit. Adobe Acrobat Pro's Redact tool does this when you click Apply, not when you just place the marks. Open-source pdf-redact-tools and several command-line utilities take a more aggressive route by re-rendering the page as an image and discarding the original stream entirely.
If a document has gone through a redaction tool, the next safety check is to run it through a metadata cleaner. Document properties do not always travel through the redaction pass. The same goes for any embedded files or annotation layers the redactor was not told about. The cleaner on the homepage will show you exactly what is left in the file's metadata before you publish it.
A safe workflow
- Apply redactions in a tool that destroys the underlying content, not one that draws shapes on top.
- Save out as a fresh PDF, not as a copy of the original with extra annotations.
- Run pdftotext (or any plain-text extractor) over the result and search for the names you redacted. If a single one comes back, the redaction failed.
- Strip the document properties and XMP metadata.
- If the stakes are high, print to PDF as a final pass: this re-renders every page from the displayed output, leaving no underlying stream to recover.
None of this is exotic, and none of it costs anything beyond a few minutes. The reason redacted documents still leak is that almost everyone treats PDFs as paper with extra steps, when they are really a small program the reader executes every time the page is shown.