A PDF in your archive is damaged. You run it through qpdf, the command finishes, and the output opens cleanly with the right page count. Repaired.
Maybe not. I damaged a 20-page document four different ways, ran three tools against each, and then checked something most repair guides skip: whether the recovered file still contains its text. One combination produced a structurally perfect PDF in which every page was blank. Correct page count, valid structure, zero content.
If you delete the original after seeing “20 pages recovered,” that is the case that costs you the document.
The four ways PDFs break
Starting from one source file (20 pages, 73,378 bytes, 58,216 characters of extractable text), I produced four corruptions that mirror what happens to archived files:
| Damage | How it happens | Does pdfinfo flag it? |
|---|---|---|
| Broken xref offset | Partial write, bad edit, disk error | Yes |
| Truncated file | Interrupted download or copy | Yes |
Missing %%EOF |
Sloppy generator, stripped trailer | No, opens fine |
| Junk before the header | Mail server preamble, concatenation | No, opens fine |
Two of the four were tolerated by every reader I tested. Before repairing anything, check whether the file is actually broken: pdfinfo flagged the first two and passed the last two.
The repair matrix, by page count
| Damage | qpdf | Ghostscript | pdftk |
|---|---|---|---|
| Broken xref | 20 pages | 20 pages | failed |
| Truncated | 20 pages | 20 pages | failed |
| Missing EOF | 20 pages | 20 pages | 20 pages |
| Junk header | 20 pages | 20 pages | 20 pages |
The obvious reading: qpdf and Ghostscript recover everything, pdftk handles the easy cases. (pdftk reported Rebuild failed: trailer not found and wrote nothing, which at least fails loudly.)
That reading is wrong.
The same matrix, by content
| Damage | qpdf | Ghostscript |
|---|---|---|
| Broken xref | 58,216 chars, full | 58,216 chars, full |
| Truncated | 0 chars, every page blank | 58,216 chars, full |
| Missing EOF | 58,216 chars, full | 58,216 chars, full |
| Junk header | 58,216 chars, full | 58,216 chars, full |
The truncated file’s qpdf output has 20 page objects and 20 /Contents references. pdfinfo reports 20 pages. It opens without complaint. But pdffonts lists no fonts, pdftotext extracts nothing, and page one renders empty. The content streams survived; the font resources they point to did not.
Ghostscript, handed the identical damaged bytes, returned the full document.
Why this happens, and why you cannot predict it
The two tools work differently. qpdf rebuilds the cross-reference table from objects it can still find and rewrites the structure around them. Ghostscript interprets the page description and generates a new PDF from what it can render. The outcome therefore depends on which objects fell on which side of the cut.
In my test document the font references live inside an object stream sitting at the 98% mark, so any truncation destroys them. qpdf recovers a page tree but no resources, which is the blank skeleton. Content streams start at the 9% mark, so Ghostscript still had something to render.
I swept the truncation depth to find where the behavior changes:
| Bytes remaining | qpdf, object-stream version | qpdf, flat version |
|---|---|---|
| 98% / 95% / 90% | 20 pages, 0 chars | 20 pages, 58,216 chars |
| 85% / 80% / 70% | 20 pages, 0 chars | 20 pages, 58,216 chars |
| 50% / 30% | 20 pages, 0 chars | 20 pages, 58,216 chars |
On this document the split is total: the object-stream layout fails silently at every depth, the flat layout survives every depth. That looks like a clean rule, and I nearly published it as one.
It is not a rule. A second document built the same way, but whose font objects happened to land near the 85% mark instead of at 17%, inverted the pattern: its flat version recovered above that point and produced the same 20-page blank below it. The variable is not object streams as such. It is whether the truncation swallowed the objects the tool needs, and that depends on where your particular producer chose to write them.
Ghostscript is not immune. It recovered content at every depth on my document, but because it regenerates rather than reconstructs, a file it cannot fully render can come back with fewer pages, or a single empty one, still exiting 0. Neither tool tells you it failed.
How to check a repair actually worked
Page count proves the page tree survived. It proves nothing about content. Three commands, all from poppler-utils:
pdfinfo repaired.pdf | grep Pages # structure survived?
pdftotext repaired.pdf - | wc -c # is there any text left?
pdffonts repaired.pdf # did the fonts come back?
Compare against the original if you still have it. An empty pdffonts table on a document that used fonts, or a character count near zero on a text document, means the repair is a shell. Scanned pages have no text to extract, so check pdfimages -list for image rows instead. The loop below tests both, so scans do not trigger it:
for f in repaired/*.pdf; do
chars=$(pdftotext "$f" - 2>/dev/null | tr -d '[:space:]' | wc -c)
imgs=$(pdfimages -list "$f" 2>/dev/null | tail -n +3 | wc -l)
if [ "$chars" -lt 20 ] && [ "$imgs" -eq 0 ]; then
echo "SUSPECT (no text, no images): $f"
fi
done
Check the page count as well. A repair that returns 1 page from a 20-page original is a failure even when text extraction looks non-empty.
What to run, in what order
| Situation | Do this |
|---|---|
| File opens and pages look right | Nothing. Missing EOF and junk headers were tolerated in every reader I tested. |
| Any real damage | Run both tools, compare content, keep the winner. This is the only approach that survived every case I measured. |
| Broken xref, file otherwise complete | qpdf alone was sufficient here: fast and lossless. |
| Truncated file | Ghostscript recovered mine at every depth and qpdf recovered none of them. Verify anyway, since the reverse is possible on a different layout. |
| Nothing produces content | Keep the damaged original. A blank repair is worse than a broken file a specialist can still work on. |
qpdf damaged.pdf fixed_qpdf.pdf
gs -sDEVICE=pdfwrite -dNOPAUSE -dQUIET -dBATCH -sOutputFile=fixed_gs.pdf damaged.pdf
pdftotext fixed_qpdf.pdf - | wc -c
pdftotext fixed_gs.pdf - | wc -c
Two lines of comparison at the end is the whole discipline. Neither tool’s exit code tells you which one worked.
On that point, if you script this: qpdf returns 3 when it emits warnings, and repairing a damaged file always emits warnings. Under set -e that reads as failure and kills a batch loop. I hit this building a batch compression script. Treat 0 and 3 as success, 2 as a real error, and judge the result by its content rather than its status.
Never repair in place
Every command here writes to a new file. qpdf --replace-input exists and is convenient, but on a file that produces a blank recovery it destroys the only copy of the damaged original, which as the truncation case shows may still hold content another tool can reach.
Test setup. Measured 26 July 2026 on Linux with qpdf 12.2.0, Ghostscript 10.05.1, pdftk-java 3.3.3 and poppler-utils. Source document: 20 pages, 73,378 bytes, 58,216 extractable characters (whitespace stripped, via pdftotext), one embedded non-base-14 font, written with object streams. Damage was applied programmatically: xref offset overwritten, file truncated by byte count, %%EOF stripped, 20 bytes prepended. The truncation sweep ran against both the object-stream version and a flat version produced with qpdf --object-streams=disable. Results depend on where a given producer writes its objects, which is why the content check matters more than any of these specific numbers.
References: qpdf manual — exit status (exit 3 means warnings, and output is still written — the reason a damaged file appears to repair cleanly) · poppler (the project behind pdfinfo, pdftotext, pdffonts and pdfimages used for the content checks)