Computers & IT

Merging PDFs: qpdf Silently Drops Every Bookmark

Merging PDFs on the command line looks like a solved problem. Pick qpdf because it is fast, or pdftk because it is familiar, and move on.

I merged five bookmarked documents four different ways and measured what came out. The standard qpdf incantation, the one in most tutorials, silently discarded all 30 bookmarks. Ghostscript kept every one of them and produced a file 56% smaller than the inputs. That is close to the opposite of what the popular guidance says.

The test

Five documents, six pages each, one embedded font, six bookmarks apiece. 50,540 bytes and 30 bookmarks going in.

Command Output Change Bookmarks Time
qpdf --empty --pages a.pdf b.pdf ... -- out.pdf 43,439 b -14.1% 0 / 30 22 ms
qpdf a.pdf --pages a.pdf b.pdf ... -- out.pdf 42,932 b -15.1% 6 / 30 18 ms
pdftk a.pdf b.pdf cat output out.pdf 48,374 b -4.3% 30 / 30 1,151 ms
gs -sDEVICE=pdfwrite -sOutputFile=out.pdf a.pdf b.pdf ... 22,290 b -55.9% 30 / 30 258 ms

All four produced 30 pages, and all four preserved the text exactly: 49,890 extractable characters in, 49,890 out, every time. No method damaged the text. What differs is structure, and structure is where the losses hide.

The bookmark result is the one to notice

qpdf --empty --pages is the form you will find in most answers, and it starts from an empty document. Document-level information such as the outline tree is taken from the primary input, and with --empty there is no primary input, so the outline has nothing to come from. Every bookmark disappears without a warning.

Using the first file as the primary input instead of --empty rescues that file’s six bookmarks and loses the other 24. It is better, and it is still a silent 80% loss.

This matters because “qpdf preserves bookmarks” appears in a lot of comparison articles. On the merge path, in the version I tested, it does not.

The reverse folklore is also stale. Ghostscript has a longstanding reputation for mangling outlines, and on 10.05.1 it kept all 30 with their destinations intact. If you last checked this years ago, check again.

Where Ghostscript’s 56% comes from

Every input used the same typeface. Count the embedded fonts in each merged file:

Merged by Embedded fonts Size
qpdf 5 43,439 b
pdftk 5 48,374 b
Ghostscript 1 22,290 b

qpdf and pdftk copy page objects and carry five copies of one font along with them. Ghostscript interprets each input and writes a single new document, so it embeds the shared font once. On documents that share fonts, that is most of the difference.

Order decides the outcome, and you cannot fix it afterwards

The obvious optimization: merge with fast qpdf, then run Ghostscript to compress. I measured it against merging with Ghostscript directly.

Pipeline Result Fonts
qpdf merge, then gs /ebook 38,723 b 5
qpdf merge, then qpdf recompress 32,657 b 5
Ghostscript merge, one pass 22,290 b 1

The two-step route lands 74% larger than the one-step route (equivalently, merging with Ghostscript directly is 42% smaller), and the Ghostscript pass does not recover the duplicate fonts at all.

The reason is more specific than it looks. I expected the five copies to survive because their subset tags differed, but they do not differ: after the Ghostscript pass all five are named FRDFEL+C059-Roman, identical byte-for-byte, and they are still five separate objects.

What decides it is the object number each font carries. Ghostscript merges font programs that arrive at the same object number and keeps the rest apart. Five inputs that each hold the font as object 9 collapse to one font; the same five inputs renumbered so the font lands at objects 14, 18, 22, 26 and 30 stay as five. Merging two copies of the identical file deduplicates, merging two files whose fonts sit at different numbers does not. A qpdf merge renumbers everything, which is exactly the condition under which Ghostscript can no longer help.

Merge order is therefore not a performance detail. It is the decision that sets your floor.

What each method quietly loses

Text survives everywhere. Other things do not, and none of the tools says so.

Feature qpdf --empty pdftk Ghostscript
Text kept kept kept
Bookmarks lost kept kept
Embedded attachments dropped kept kept
Form fields (AcroForm) kept kept destroyed
Encryption preserved preserved stripped

Two of these are worth spelling out because they are silent and they matter.

Ghostscript destroys form fields. A one-field document with the value hello came back from a Ghostscript pass with zero fields, no /AcroForm entry and no /Widget object anywhere in the file. The page still renders; the form is gone. This is not merge-specific, it happens on any single-file pass too.

Ghostscript strips encryption. An AES-256 encrypted input came out unencrypted and opened with no password at all. If you are merging protected documents, the output is not protected.

On the other side, qpdf --empty drops embedded file attachments along with the bookmarks, for the same reason: with no primary input there is no document-level information to carry forward. qpdf a.pdf --pages ... keeps them.

What each tool is actually for

Goal Use Cost
Smallest output, bookmarks kept Ghostscript Slower than qpdf, and it rewrites the file: form fields and encryption do not survive
Fastest possible merge, page content untouched qpdf Bookmarks lost, fonts duplicated, attachments dropped by --empty
Bookmarks kept without rewriting content pdftk Slowest here, no size benefit
Documents with nothing in common Ghostscript, still Gap narrows but does not reverse: 19.9% smaller vs qpdf 13.5%, and still the only fast option keeping bookmarks
gs -sDEVICE=pdfwrite -dNOPAUSE -dQUIET -dBATCH -sOutputFile=merged.pdf a.pdf b.pdf c.pdf
qpdf --empty --pages a.pdf b.pdf c.pdf -- merged.pdf
pdftk a.pdf b.pdf c.pdf cat output merged.pdf

The size advantage depends on inputs sharing resources, but less than I expected. Five documents built from one template is the good case. I also merged five documents using five different typefaces, sharing nothing: Ghostscript still came out smallest at 19.9% below the inputs against qpdf’s 13.5%, and still kept all 30 bookmarks. The gap narrows; it does not reverse.

Check the merge, not just the page count

Page count is the one thing every method got right, which makes it useless as a check. Verify the parts that actually varied:

pdfinfo merged.pdf | grep Pages
pdftk merged.pdf dump_data | grep -c BookmarkTitle
pdffonts merged.pdf | tail -n +3 | wc -l
pdftotext merged.pdf - | wc -c

Bookmark count against the sum of the inputs, font count against how many distinct typefaces you expect, character count against the originals. A merge that keeps 30 pages and drops 30 bookmarks looks perfect until someone opens the outline panel.

One scripting note: qpdf returns exit code 3 for warnings, and warnings are routine. Under set -e that reads as failure, which is how it can kill a batch loop, as I found while building a batch compression script. Treat 0 and 3 as success and judge by the output.

If a merge fails outright, the inputs may already be damaged. Repairing them first has its own trap: a repaired file can report the right page count and contain nothing.


Test setup. Measured 26 July 2026 on Linux with qpdf 12.2.0, Ghostscript 10.05.1, pdftk-java 3.3.3 and poppler-utils. Inputs: five PDFs of six pages each, generated from PostScript with pdfmark outlines, one embedded non-base-14 font shared across all five, 50,540 bytes and 30 bookmarks total. Bookmark counts from pdftk dump_data, font counts from pdffonts, text from pdftotext with whitespace stripped. Timings are single runs on one core. Results depend on how much your inputs share: documents with no common resources will not show the same size gap.

References: qpdf manual (--pages merge behaviour — the documentation does not claim outline preservation, which matches what I measured) · Ghostscript High Level Devices (pdfmark handling, including outlines)