Computers & IT

Why Ghostscript Made Your PDF Bigger (And How to Fix It)

You ran a PDF through Ghostscript to shrink it. The output came back bigger than the input — sometimes by 30% or more. No errors, no warnings, just a larger file.

Trying a more aggressive -dPDFSETTINGS preset won’t help, and neither will the advice you’ll find in most older threads. Here is what actually causes it, measured on Ghostscript 10.05.1.

First, correct a widely repeated claim

Search this problem and you will quickly find the explanation that Ghostscript’s pdfwrite device cannot write object streams (ObjStm) or cross-reference streams (XRefStm), so any file using them necessarily inflates. That explanation is well sourced — it comes from Ghostscript’s own blog — but it is out of date.

Since Ghostscript 10.03 (2024), pdfwrite writes both by default. The parameters -dWriteObjStms and -dWriteXRefStm both default to true. Checking a plain pdfwrite output on 10.05.1:

$ grep -ac '^xref$' output.pdf
0                                  # no plain xref table at all
$ qpdf --show-object=trailer output.pdf | grep -o '/Type */XRef'
/Type /XRef                        # cross-reference stream present

And those structures are doing real work. Disabling them on the same input:

Command Output
pdfwrite defaults 24,838 bytes
-dWriteObjStms=false -dWriteXRefStm=false 29,330 bytes

Ghostscript’s own structural compression saved 4,492 bytes here. So “pdfwrite can’t compress structure” is not the reason your file grew.

One caveat worth knowing: both switches are forced off when the output is below PDF 1.5. If your workflow passes -dCompatibilityLevel=1.4, you silently get the old, larger structure back — and in that case the outdated advice does apply again.

The real cause: pdfwrite leaves page dictionaries outside its object streams

The page content is not the problem. Comparing the compressed content streams before and after Ghostscript on my test file:

Source:            41 content streams, 11,370 bytes total
After Ghostscript: 41 content streams, 11,370 bytes total   # byte-for-byte identical

Nothing was re-encoded and nothing got more verbose. What changed is where the structural objects live.

File Size Packed into ObjStm /Page dicts left uncompressed
Source (written by qpdf) 18,643 bytes 84 0
After Ghostscript 24,838 bytes 82 40 — 5,526 bytes

qpdf packs all 40 page dictionaries into its object stream, where they get Flate-compressed. Ghostscript writes them as plain, uncompressed top-level objects:

4 0 obj
<</Type/Page/MediaBox [0 0 595 842]
/Rotate 0/Parent 3 0 R
/Resources<</ProcSet[/PDF /Text]
/Font 9 0 R
>>
/Contents 5 0 R
>>
endobj                          # one of these per page, uncompressed

Those 5,526 bytes account for nearly all of the 6,195-byte increase. Ghostscript’s object stream is working — it’s simply carrying less.

The giveaway: the overhead scales with page count, not content. Running the same comparison on a 600-page document:

Pages Source After Ghostscript Overhead per page
40 18,643 bytes 24,838 bytes 154.9 bytes
600 106,389 bytes 200,489 bytes 156.8 bytes

A near-constant cost per page — roughly the size of one uncompressed page dictionary plus its cross-reference entry. If the cause were verbose content, the penalty would track text volume instead.

Where page dictionaries are storedqpdf packs all 40 page dictionaries into a compressed object stream, producing an 18,643 byte file. Ghostscript writes the same 40 page dictionaries as uncompressed top-level objects totalling 5,526 bytes, producing a 24,838 byte file. The overhead is about 155 bytes per page regardless of document length.Same 40-page document · where the 40 page dictionaries end upqpdfall 40 packed into the object streamObjStm — Flate compresseduncompressed page dicts: 018,643 bytesGhostscriptpage dicts left outside, uncompressedObjStm — carries less…40uncompressed page dicts: 40 — 5,526 B24,838 bytes (+33%)The overhead tracks page count, not content40 pages18,643 → 24,838 · 154.9 B/page600 pages106,389 → 200,489 · 156.8 B/page
Where page dictionaries are stored

Measured: same command, opposite outcomes

To confirm the source structure is what decides the outcome, I built two copies of one document differing only in whether object streams are used:

qpdf --object-streams=generate base.pdf objstm.pdf     # 18,643 bytes
qpdf --object-streams=disable  base.pdf noobjstm.pdf   # 28,125 bytes

Source written compactly (18,643 bytes)

Command Output Change
-dPDFSETTINGS=/ebook 24,514 bytes +31.5%
-dPDFSETTINGS=/screen 24,861 bytes +33.4%
-dPDFSETTINGS=/prepress 27,435 bytes +47.2%
-dPDFSETTINGS=/printer 31,266 bytes +67.7%
no preset 24,838 bytes +33.2%

Same document written loosely (28,125 bytes)

Command Output Change
Ghostscript, no preset 24,838 bytes −11.7%

Both runs produced the same 24,838-byte file. Ghostscript writes its own structure regardless of the input, so whether that reads as “compression” or “bloat” depends entirely on how efficiently the source was written.

Note that /screen, the most aggressive preset, still produced a larger file. This document had no images, so downsampling had nothing to act on while the rewrite cost applied anyway.

The same file through qpdf

Command Output Change
qpdf --object-streams=preserve 18,643 bytes ±0%
qpdf --object-streams=generate --recompress-flate --compression-level=9 18,530 bytes −0.6%

qpdf restructures without re-interpreting page content, so it can’t inflate the way pdfwrite does. For this test file — text-heavy, no images, and using the non-embedded base-14 fonts — near-zero change is the correct answer: there was nothing left to squeeze. That does not generalise to every text PDF: if the fonts are embedded in full, a great deal is still left to squeeze, and qpdf cannot reach it, because qpdf never subsets fonts.

Check what your file actually uses

Ask qpdf what’s in the trailer:

qpdf --show-object=trailer yourfile.pdf | grep -o '/Type */XRef'

/Type /XRef means the file uses a cross-reference stream, and almost always object streams with it. No output means a plain xref table. For an exact count of object stream containers:

qpdf --json=2 yourfile.pdf | grep -c '"/Type": "/ObjStm"'

Don’t grep the raw bytes for ObjStm. Ghostscript writes its own command line into the output as a %%Invocation: comment, so a file produced with -dWriteObjStms=false — containing no object streams whatsoever — still matches a naive grep. Producer strings, XMP metadata and document text cause the same false positive.

The PDF version header tells you nothing either. All of my test files report %PDF-1.7 whether or not they use object streams. Version 1.5+ means the feature is permitted, not that it is used.

What to do instead

Your situation Best tool
Text-heavy PDF, fonts not embedded (base-14) qpdf — restructures without re-interpreting content
Text-heavy PDF, fonts embedded but not subset (pdffonts shows emb yes / sub no) Ghostscript — it subsets the fonts; measured −66.5% where qpdf gave −0.0%
Scanned pages or large images Ghostscript — image downsampling is what it’s for
Images in an already-compact file Ghostscript alone — the qpdf follow-up gains under 0.2%
You just need it smaller, any means Run both, keep the smaller result

That third row deserves a correction. I originally suggested chaining the two — Ghostscript first for image downsampling, then qpdf to recompress the structure Ghostscript rebuilt. The logic is sound, so I measured it:

File After Ghostscript After qpdf pass Extra gain
300dpi scan 270,988 b 270,787 b 201 bytes (0.07%)
150dpi scan 285,462 b 284,924 b 538 bytes (0.19%)

Under a fifth of one percent. On an image-heavy file the structural overhead is a rounding error next to the image data, and on a text-heavy file whose fonts are not embedded — where structure is all that is left — Ghostscript has nothing to work with. The exception is a text file carrying fully embedded, non-subset fonts: there Ghostscript subsets them and wins outright (I measured 76,447 b to 25,572 b, −66.5% under /ebook, against −0.0% for qpdf). Skip the second pass and use the right single tool.

Quick reference

Symptom Likely cause Fix
Output larger, no images in file Page dictionaries written uncompressed Check pdffonts first: if no font is embedded, use qpdf instead; if a font shows emb yes / sub no, keep Ghostscript — subsetting outweighs the ~155 b/page cost
Output larger even with /screen Presets affect images, not structure — and no font was embedded for subsetting to shrink Use qpdf instead
Output identical size, has images Downsample threshold not met Lower the preset or the threshold
Output smaller but blurry Preset target far below source DPI Move up one preset

If your file does contain images and simply refuses to shrink, that’s a different mechanism — Ghostscript only downsamples when the source resolution exceeds 1.5× the preset target. I measured that separately in Ghostscript /ebook not reducing PDF size.


Test setup. Measured 22 July 2026 · Ghostscript 10.05.1 · qpdf 12.2.0 · Linux. The test document was a 40-page text-only PDF generated from PostScript, then written twice by qpdf — once with --object-streams=generate, once with --object-streams=disable — so the two inputs differ only in structural compression. Object counts come from the /Size entry and the /N value of each ObjStm. Figures are actual output file sizes; your numbers will differ with content and Ghostscript version. Note in particular that behavior changed in Ghostscript 10.03 — measurements taken on 9.x will not match.

References: Ghostscript High Level Devices — WriteObjStms, WriteXRefStm · Ghostscript 10.03.0 release notes · qpdf manual — object stream options