You extract five pages from a sixty-page report. The result should be about a twelfth of the original. Instead it comes back at 91% of the full file.
That is not a rounding error and it is not your imagination. I measured it across three tools and four extraction sizes, and the cause turns out to have almost nothing to do with the pages you asked for.
The measurements
Source: a 60-page report, 339,644 bytes, five embedded typefaces.
| Pages extracted | qpdf | pdftk | Ghostscript | qpdf vs gs |
|---|---|---|---|---|
| 1 of 60 | 63,835 b | 62,726 b | 6,946 b | 9.2x |
| 5 of 60 | 309,978 b | 311,702 b | 23,129 b | 13.4x |
| 15 of 60 | 315,301 b | 318,036 b | 28,462 b | 11.1x |
| 30 of 60 | 323,324 b | 327,612 b | 36,530 b | 8.9x |
Look at the qpdf column going down. Five pages costs 310 KB. Thirty pages costs 323 KB. Six times as many pages, four percent more file. Whatever is filling those files, it is not the page content.
What is actually in there
Count the fonts instead of the pages:
| Pages extracted | Fonts in the qpdf output | Size |
|---|---|---|
| 1 | 1 | 63,835 b |
| 5 | 5 | 309,978 b |
| 15 | 5 | 315,301 b |
| 30 | 5 | 323,324 b |
In this document the file size tracks the font count, not the page count, because five unsubsetted faces weigh about 300 KB while a page of text weighs a few hundred bytes. My source cycles through five typefaces, so page five is where the extract has touched all of them. After that the fonts are already paid for and additional pages cost only their text.
Invert that ratio and the rule inverts with it. On an image-heavy report with a single embedded face, the same extractions ran 76 KB, 127 KB, 254 KB and 445 KB, scaling with pages exactly as you would expect. What holds generally is narrower: an unsubsetted font is a fixed toll paid the first time the range touches it, and everything else still scales with pages.
pdffonts on the five-page extracts shows the difference plainly:
| Extracted by | Font entries |
|---|---|
| qpdf | C059-Roman, P052-Roman, … all emb=yes sub=no |
| Ghostscript | QSHGYR+C059-Roman, PNXBHT+P052-Roman, … all emb=yes sub=yes |
qpdf and pdftk copy the font programs across unchanged. If the source embedded a complete typeface, your five-page extract carries that complete typeface, every glyph of it, for the handful of characters those five pages use. Ghostscript regenerates the document and subsets each face down to the glyphs that survive, which is where the 13x comes from.
The flag that gets recommended, and when it actually helps
Search this problem and you will be pointed at --remove-unreferenced-resources. It sounds exactly right: strip the resources the extracted pages no longer reference. I ran all three settings on the one-page extract:
| Setting | Output |
|---|---|
auto (default) |
63,835 b |
yes |
63,833 b |
no |
63,835 b |
Two bytes, because this source gives every page its own resource dictionary holding just the one face that page uses. There is nothing unreferenced left to drop.
That is not universal, and assuming it is will cost you. I rebuilt the same document with a single resource dictionary shared across all sixty pages, the way many LaTeX and word-processor exports emit it, then extracted one page again:
| Setting | Output | Fonts |
|---|---|---|
auto (default) |
310,669 b | 5 |
yes |
64,169 b | 1 |
no |
310,669 b | 5 |
A 4.8x difference, with identical extracted text in all three. auto declines to prune a dictionary that other pages share, which is the safe call and also the expensive one. So run the default, look at the result, and try =yes before concluding the flag is useless.
What it can never do is shrink a font the page genuinely uses. On this source one face is 73,884 bytes where the glyphs actually on the page need 3,188: about twenty times larger than it has to be.
When this does not happen
The effect needs one specific condition, and it is worth checking before you reach for a heavier tool. I rebuilt the same 60-page source with its fonts already subset, then extracted one page again:
| Source fonts | qpdf, 1 page | Ghostscript, 1 page |
|---|---|---|
| Embedded, not subset | 63,835 b | 6,946 b |
| Already subset | 6,895 b | 6,946 b |
With subset fonts the gap collapses, from 9.2x to roughly 1.1x here. What remains depends on how much narrower one page’s glyph set is than the whole document’s: qpdf carries the document-wide subset, Ghostscript re-cuts it to the extracted page, so on a source whose pages differ sharply in the characters they use the residue is still worth measuring. Either way the rule is not “use Ghostscript to extract pages.” It is that unsubsetted source fonts make a copy-based extraction cost one full font program per distinct face the range touches, however few pages you asked for, and only a tool that re-subsets can fix it.
Check first, then choose
pdffonts source.pdf
Read the sub column. If embedded fonts show no, a copy-based extract will carry them whole and Ghostscript is worth the slower pass. If they already show yes, use qpdf and keep the speed.
| Situation | Use |
|---|---|
| Source fonts unsubset, size matters | Ghostscript, up to 13x smaller here |
| Source fonts already subset | qpdf, same size and much faster |
| You must not re-encode anything | qpdf, and accept the size |
| Bookmarks matter | Verify afterwards. On a 60-bookmark source, extracting five pages left qpdf with all 60 (55 pointing at pages that no longer exist), pdftk with none at all, and Ghostscript with the correct 5 |
| Form fields or encryption involved | Not Ghostscript, it destroys both |
qpdf source.pdf --pages . 1-5 -- out.pdf
pdftk source.pdf cat 1-5 output out.pdf
gs -sDEVICE=pdfwrite -dNOPAUSE -dQUIET -dBATCH -dFirstPage=1 -dLastPage=5 -sOutputFile=out.pdf source.pdf
That last row matters more than the size question if your document is a form or is encrypted. Ghostscript rewrites the file, and in doing so it removes AcroForm fields entirely and outputs an unencrypted result. I measured both while comparing merge tools, and the same applies to extraction.
-dFirstPage and -dLastPage take a contiguous range, but Ghostscript is not limited to one. For an arbitrary selection use -sPageList, which accepts a comma-separated list and ranges together:
gs -sDEVICE=pdfwrite -dNOPAUSE -dQUIET -dBATCH -sPageList=1,5,9-12 -sOutputFile=out.pdf source.pdf
Verify the extract
pdfinfo out.pdf | grep Pages
pdffonts out.pdf | tail -n +3 | wc -l
pdftotext out.pdf - | wc -c
pdftk out.pdf dump_data | grep -c BookmarkTitle
Page count confirms you got the range you asked for. Font count tells you whether the typefaces came along whole. Bookmark count catches the outline damage above. Character count catches the case where a rewrite dropped content, which is a real risk with damaged sources: a rebuilt file can report the right page count and contain nothing at all.
Test setup. Measured 27 July 2026 on Linux with qpdf 12.2.0, Ghostscript 10.05.1, pdftk-java 3.3.3 and poppler-utils. Source: a 60-page PostScript-generated report, 339,644 bytes, cycling through five URW typefaces embedded without subsetting, plus a subset-font rebuild of the same document for the control. Sizes are actual output files. How much you gain depends on how many distinct unsubsetted faces your source embeds: one face means one copy to shrink, five means five.
References: qpdf manual — page selection (the --pages syntax used above) · Ghostscript High Level Devices (SubsetFonts, which is why the extracted file keeps the whole typeface unless Ghostscript re-writes it)