Computers & IT

Extracting 5 Pages From a 60-Page PDF Gave Me 91% of the File

You extract five pages from a sixty-page report. The result should be about a twelfth of the original. Instead it comes back at 91% of the full file.

That is not a rounding error and it is not your imagination. I measured it across three tools and four extraction sizes, and the cause turns out to have almost nothing to do with the pages you asked for.

The measurements

Source: a 60-page report, 339,644 bytes, five embedded typefaces.

Pages extracted qpdf pdftk Ghostscript qpdf vs gs
1 of 60 63,835 b 62,726 b 6,946 b 9.2x
5 of 60 309,978 b 311,702 b 23,129 b 13.4x
15 of 60 315,301 b 318,036 b 28,462 b 11.1x
30 of 60 323,324 b 327,612 b 36,530 b 8.9x

Look at the qpdf column going down. Five pages costs 310 KB. Thirty pages costs 323 KB. Six times as many pages, four percent more file. Whatever is filling those files, it is not the page content.

What is actually in there

Count the fonts instead of the pages:

Pages extracted Fonts in the qpdf output Size
1 1 63,835 b
5 5 309,978 b
15 5 315,301 b
30 5 323,324 b

In this document the file size tracks the font count, not the page count, because five unsubsetted faces weigh about 300 KB while a page of text weighs a few hundred bytes. My source cycles through five typefaces, so page five is where the extract has touched all of them. After that the fonts are already paid for and additional pages cost only their text.

Invert that ratio and the rule inverts with it. On an image-heavy report with a single embedded face, the same extractions ran 76 KB, 127 KB, 254 KB and 445 KB, scaling with pages exactly as you would expect. What holds generally is narrower: an unsubsetted font is a fixed toll paid the first time the range touches it, and everything else still scales with pages.

pdffonts on the five-page extracts shows the difference plainly:

Extracted by Font entries
qpdf C059-Roman, P052-Roman, … all emb=yes sub=no
Ghostscript QSHGYR+C059-Roman, PNXBHT+P052-Roman, … all emb=yes sub=yes

qpdf and pdftk copy the font programs across unchanged. If the source embedded a complete typeface, your five-page extract carries that complete typeface, every glyph of it, for the handful of characters those five pages use. Ghostscript regenerates the document and subsets each face down to the glyphs that survive, which is where the 13x comes from.

The flag that gets recommended, and when it actually helps

Search this problem and you will be pointed at --remove-unreferenced-resources. It sounds exactly right: strip the resources the extracted pages no longer reference. I ran all three settings on the one-page extract:

Setting Output
auto (default) 63,835 b
yes 63,833 b
no 63,835 b

Two bytes, because this source gives every page its own resource dictionary holding just the one face that page uses. There is nothing unreferenced left to drop.

That is not universal, and assuming it is will cost you. I rebuilt the same document with a single resource dictionary shared across all sixty pages, the way many LaTeX and word-processor exports emit it, then extracted one page again:

Setting Output Fonts
auto (default) 310,669 b 5
yes 64,169 b 1
no 310,669 b 5

A 4.8x difference, with identical extracted text in all three. auto declines to prune a dictionary that other pages share, which is the safe call and also the expensive one. So run the default, look at the result, and try =yes before concluding the flag is useless.

What it can never do is shrink a font the page genuinely uses. On this source one face is 73,884 bytes where the glyphs actually on the page need 3,188: about twenty times larger than it has to be.

When this does not happen

The effect needs one specific condition, and it is worth checking before you reach for a heavier tool. I rebuilt the same 60-page source with its fonts already subset, then extracted one page again:

Source fonts qpdf, 1 page Ghostscript, 1 page
Embedded, not subset 63,835 b 6,946 b
Already subset 6,895 b 6,946 b

With subset fonts the gap collapses, from 9.2x to roughly 1.1x here. What remains depends on how much narrower one page’s glyph set is than the whole document’s: qpdf carries the document-wide subset, Ghostscript re-cuts it to the extracted page, so on a source whose pages differ sharply in the characters they use the residue is still worth measuring. Either way the rule is not “use Ghostscript to extract pages.” It is that unsubsetted source fonts make a copy-based extraction cost one full font program per distinct face the range touches, however few pages you asked for, and only a tool that re-subsets can fix it.

Check first, then choose

pdffonts source.pdf

Read the sub column. If embedded fonts show no, a copy-based extract will carry them whole and Ghostscript is worth the slower pass. If they already show yes, use qpdf and keep the speed.

Situation Use
Source fonts unsubset, size matters Ghostscript, up to 13x smaller here
Source fonts already subset qpdf, same size and much faster
You must not re-encode anything qpdf, and accept the size
Bookmarks matter Verify afterwards. On a 60-bookmark source, extracting five pages left qpdf with all 60 (55 pointing at pages that no longer exist), pdftk with none at all, and Ghostscript with the correct 5
Form fields or encryption involved Not Ghostscript, it destroys both
qpdf source.pdf --pages . 1-5 -- out.pdf
pdftk source.pdf cat 1-5 output out.pdf
gs -sDEVICE=pdfwrite -dNOPAUSE -dQUIET -dBATCH -dFirstPage=1 -dLastPage=5 -sOutputFile=out.pdf source.pdf

That last row matters more than the size question if your document is a form or is encrypted. Ghostscript rewrites the file, and in doing so it removes AcroForm fields entirely and outputs an unencrypted result. I measured both while comparing merge tools, and the same applies to extraction.

-dFirstPage and -dLastPage take a contiguous range, but Ghostscript is not limited to one. For an arbitrary selection use -sPageList, which accepts a comma-separated list and ranges together:

gs -sDEVICE=pdfwrite -dNOPAUSE -dQUIET -dBATCH -sPageList=1,5,9-12 -sOutputFile=out.pdf source.pdf

Verify the extract

pdfinfo out.pdf | grep Pages
pdffonts out.pdf | tail -n +3 | wc -l
pdftotext out.pdf - | wc -c
pdftk out.pdf dump_data | grep -c BookmarkTitle

Page count confirms you got the range you asked for. Font count tells you whether the typefaces came along whole. Bookmark count catches the outline damage above. Character count catches the case where a rewrite dropped content, which is a real risk with damaged sources: a rebuilt file can report the right page count and contain nothing at all.


Test setup. Measured 27 July 2026 on Linux with qpdf 12.2.0, Ghostscript 10.05.1, pdftk-java 3.3.3 and poppler-utils. Source: a 60-page PostScript-generated report, 339,644 bytes, cycling through five URW typefaces embedded without subsetting, plus a subset-font rebuild of the same document for the control. Sizes are actual output files. How much you gain depends on how many distinct unsubsetted faces your source embeds: one face means one copy to shrink, five means five.

References: qpdf manual — page selection (the --pages syntax used above) · Ghostscript High Level Devices (SubsetFonts, which is why the extracted file keeps the whole typeface unless Ghostscript re-writes it)