Why PDF redaction fails

The failure is always the same shape, and it takes ten seconds to check for.

A PDF is not a picture of a page. It is a set of instructions: draw this text, in this font, at these coordinates. When you draw a black rectangle over a name in most PDF editors, you have added one more instruction — draw a filled black box at these coordinates. You have not touched the instruction that draws the name.

The name is still in the file. It is underneath, in the same place it always was, still selectable, still searchable, still extractable by anything that reads the file rather than looks at it. What you have done is closer to laying a piece of tape over a screen than to deleting anything.

The ten-second test

Open the finished PDF, select the blacked-out region with your cursor, copy it, and paste into any text editor. If the hidden text appears, your redaction is decorative. This is the entire attack — no tools, no expertise, no intent required. A curious reader finds it by accident.

Two further checks are worth doing on anything sensitive: use your PDF reader's search function to search for the name you removed, and open the file's document properties to see what the metadata still says.

It happens to people who should know better

Paul Manafort's legal team, January 2019

Manafort's defence lawyers filed a court brief with sections blacked out. Journalists copied the blacked-out passages and pasted them into a word processor within hours, revealing material the filing was specifically intended to keep sealed — including that Manafort had shared Trump campaign polling data with an associate linked to Russian intelligence. The filing was pulled and replaced with a properly redacted version, long after the contents had been reported worldwide.

These were experienced federal litigators, filing in a case under intense scrutiny, using professional software. The technique that defeated them was highlight, copy, paste.

The TSA screening manual, December 2009

The US Transportation Security Administration published its Standard Operating Procedures for airport screening on a public government website, with sensitive passages covered by black boxes. The same copy-and-paste recovered them. The exposed material described screening procedures across 450 US airports, which bags get checked for explosives and how often, handling of intelligence assets in transit, and images of official identification cards.

Why the good intention still produces the bad file

Nobody in either case was careless in the ordinary sense. They drew boxes over the right words. The problem is that the interface tells you nothing is wrong: on screen, a covered name looks exactly like a removed name. There is no warning, no visual difference, and no moment where the software says “this text is still in the document”. You find out when someone else does.

This is why redaction is one of the few tasks where you should not trust what you can see. The check has to be mechanical.

What an actual removal looks like

There are only a few ways to genuinely destroy the text, and they all amount to one idea: the text must stop existing in the file.

  • Rasterise the page. Convert the finished page to an image, so the file contains pixels rather than characters. Nothing survives to be selected, because there is no text layer left anywhere on the page — not under the boxes, not beside them.
  • Use a true redaction tool that deletes the underlying content stream. Acrobat Pro's redaction feature does this properly if you use the redaction tool specifically and apply it — as distinct from the drawing or markup tools, which do not.
  • Print to paper, black it out with a marker, and rescan. Crude, and it still leaves the scan open to OCR of anything the marker missed, but the digital text layer is genuinely gone.

What does not work: drawing shapes, highlighting in black, changing the text colour to match the background, or covering text with a white box. Every one of those leaves the original characters in the file.

Metadata is the second leak

Even a correctly redacted page can betray a document through what sits around it. PDFs commonly carry the author's name, the originating software, the file path it was saved from, and creation and modification timestamps. A file path alone can name a client, a matter number, or an employee. Check document properties before you send anything, no matter how clean the pages look.

How BLACKOUT handles this

BLACKOUT takes the rasterising approach: every page of an export is flattened to an image, so no text layer survives anywhere in the finished file — not beneath a box and not elsewhere on the page. The copy-paste test returns nothing because there is nothing to return.

The trade-off is honest and worth stating: because the page becomes an image, the exported file is not searchable or selectable, and it is larger than the original. For a document you are about to disclose, that is usually the right trade. For a working draft you still need to search, it is not.

The whole process runs in your browser. The document is never uploaded, which also means it is never sitting on someone else's server waiting to be redacted. Why that matters for confidential files.

Sources

← All guides