Computers & IT

Your PDF Repair Probably Worked. Probably.

A PDF in your archive is damaged. You run it through qpdf, the command finishes, and the output opens cleanly with the right page count. Repaired.

Maybe not. I damaged a 20-page document four different ways, ran three tools against each, and then checked something most repair guides skip: whether the recovered file still contains its text. One combination produced a structurally perfect PDF in which every page was blank. Correct page count, valid structure, zero content.

If you delete the original after seeing “20 pages recovered,” that is the case that costs you the document.

The four ways PDFs break

Starting from one source file (20 pages, 73,378 bytes, 58,216 characters of extractable text), I produced four corruptions that mirror what happens to archived files:

Damage How it happens Does pdfinfo flag it?
Broken xref offset Partial write, bad edit, disk error Yes
Truncated file Interrupted download or copy Yes
Missing %%EOF Sloppy generator, stripped trailer No, opens fine
Junk before the header Mail server preamble, concatenation No, opens fine

Two of the four were tolerated by every reader I tested. Before repairing anything, check whether the file is actually broken: pdfinfo flagged the first two and passed the last two.

The repair matrix, by page count

Damage qpdf Ghostscript pdftk
Broken xref 20 pages 20 pages failed
Truncated 20 pages 20 pages failed
Missing EOF 20 pages 20 pages 20 pages
Junk header 20 pages 20 pages 20 pages

The obvious reading: qpdf and Ghostscript recover everything, pdftk handles the easy cases. (pdftk reported Rebuild failed: trailer not found and wrote nothing, which at least fails loudly.)

That reading is wrong.

The same matrix, by content

Damage qpdf Ghostscript
Broken xref 58,216 chars, full 58,216 chars, full
Truncated 0 chars, every page blank 58,216 chars, full
Missing EOF 58,216 chars, full 58,216 chars, full
Junk header 58,216 chars, full 58,216 chars, full

The truncated file’s qpdf output has 20 page objects and 20 /Contents references. pdfinfo reports 20 pages. It opens without complaint. But pdffonts lists no fonts, pdftotext extracts nothing, and page one renders empty. The content streams survived; the font resources they point to did not.

Ghostscript, handed the identical damaged bytes, returned the full document.

Why this happens, and why you cannot predict it

The two tools work differently. qpdf rebuilds the cross-reference table from objects it can still find and rewrites the structure around them. Ghostscript interprets the page description and generates a new PDF from what it can render. The outcome therefore depends on which objects fell on which side of the cut.

In my test document the font references live inside an object stream sitting at the 98% mark, so any truncation destroys them. qpdf recovers a page tree but no resources, which is the blank skeleton. Content streams start at the 9% mark, so Ghostscript still had something to render.

I swept the truncation depth to find where the behavior changes:

Bytes remaining qpdf, object-stream version qpdf, flat version
98% / 95% / 90% 20 pages, 0 chars 20 pages, 58,216 chars
85% / 80% / 70% 20 pages, 0 chars 20 pages, 58,216 chars
50% / 30% 20 pages, 0 chars 20 pages, 58,216 chars

On this document the split is total: the object-stream layout fails silently at every depth, the flat layout survives every depth. That looks like a clean rule, and I nearly published it as one.

It is not a rule. A second document built the same way, but whose font objects happened to land near the 85% mark instead of at 17%, inverted the pattern: its flat version recovered above that point and produced the same 20-page blank below it. The variable is not object streams as such. It is whether the truncation swallowed the objects the tool needs, and that depends on where your particular producer chose to write them.

Ghostscript is not immune. It recovered content at every depth on my document, but because it regenerates rather than reconstructs, a file it cannot fully render can come back with fewer pages, or a single empty one, still exiting 0. Neither tool tells you it failed.

How to check a repair actually worked

Page count proves the page tree survived. It proves nothing about content. Three commands, all from poppler-utils:

pdfinfo repaired.pdf | grep Pages      # structure survived?
pdftotext repaired.pdf - | wc -c       # is there any text left?
pdffonts repaired.pdf                  # did the fonts come back?

Compare against the original if you still have it. An empty pdffonts table on a document that used fonts, or a character count near zero on a text document, means the repair is a shell. Scanned pages have no text to extract, so check pdfimages -list for image rows instead. The loop below tests both, so scans do not trigger it:

for f in repaired/*.pdf; do
  chars=$(pdftotext "$f" - 2>/dev/null | tr -d '[:space:]' | wc -c)
  imgs=$(pdfimages -list "$f" 2>/dev/null | tail -n +3 | wc -l)
  if [ "$chars" -lt 20 ] && [ "$imgs" -eq 0 ]; then
    echo "SUSPECT (no text, no images): $f"
  fi
done

Check the page count as well. A repair that returns 1 page from a 20-page original is a failure even when text extraction looks non-empty.

What to run, in what order

Situation Do this
File opens and pages look right Nothing. Missing EOF and junk headers were tolerated in every reader I tested.
Any real damage Run both tools, compare content, keep the winner. This is the only approach that survived every case I measured.
Broken xref, file otherwise complete qpdf alone was sufficient here: fast and lossless.
Truncated file Ghostscript recovered mine at every depth and qpdf recovered none of them. Verify anyway, since the reverse is possible on a different layout.
Nothing produces content Keep the damaged original. A blank repair is worse than a broken file a specialist can still work on.
qpdf damaged.pdf fixed_qpdf.pdf
gs -sDEVICE=pdfwrite -dNOPAUSE -dQUIET -dBATCH -sOutputFile=fixed_gs.pdf damaged.pdf
pdftotext fixed_qpdf.pdf - | wc -c
pdftotext fixed_gs.pdf - | wc -c

Two lines of comparison at the end is the whole discipline. Neither tool’s exit code tells you which one worked.

On that point, if you script this: qpdf returns 3 when it emits warnings, and repairing a damaged file always emits warnings. Under set -e that reads as failure and kills a batch loop. I hit this building a batch compression script. Treat 0 and 3 as success, 2 as a real error, and judge the result by its content rather than its status.

Never repair in place

Every command here writes to a new file. qpdf --replace-input exists and is convenient, but on a file that produces a blank recovery it destroys the only copy of the damaged original, which as the truncation case shows may still hold content another tool can reach.


Test setup. Measured 26 July 2026 on Linux with qpdf 12.2.0, Ghostscript 10.05.1, pdftk-java 3.3.3 and poppler-utils. Source document: 20 pages, 73,378 bytes, 58,216 extractable characters (whitespace stripped, via pdftotext), one embedded non-base-14 font, written with object streams. Damage was applied programmatically: xref offset overwritten, file truncated by byte count, %%EOF stripped, 20 bytes prepended. The truncation sweep ran against both the object-stream version and a flat version produced with qpdf --object-streams=disable. Results depend on where a given producer writes its objects, which is why the content check matters more than any of these specific numbers.

References: qpdf manual — exit status (exit 3 means warnings, and output is still written — the reason a damaged file appears to repair cleanly) · poppler (the project behind pdfinfo, pdftotext, pdffonts and pdfimages used for the content checks)