Computers & IT

Batch PDF Compression: Why the One-Line Loop Wastes 60% of Your Savings

Search for batch PDF compression and you will find the same snippet everywhere:

for f in *.pdf; do gs -sDEVICE=pdfwrite -dPDFSETTINGS=/ebook -dNOPAUSE -dQUIET -dBATCH -sOutputFile="out/$f" "$f"; done

It works, in the sense that it terminates. On a realistic mixed folder it also throws away most of the compression you could have had, and takes longer doing it.

I built a 40-file corpus that looks like a real document archive — scanned pages, text documents with embedded fonts, text documents without — and measured three strategies against it.

The corpus

Type Count
Text, fonts embedded but not subset 15
Text, no embedded fonts 15
Scanned pages, 150dpi 7
Scanned pages, 300dpi 3
Total 40 files, 19.18 MB

That mix matters. A folder of nothing but scans makes any Ghostscript loop look brilliant; a folder of nothing but text makes it look broken. Real archives contain both.

The three strategies

Strategy Output Reduction Time
① Ghostscript /ebook on everything 14.16 MB −26.2% 11.99 s
② Route by images only 3.48 MB −81.8% 5.03 s
③ Route by images and fonts 2.73 MB −85.8% 9.90 s

The naive loop left 11.4 MB on the table — a result five times larger than the best strategy, while taking longer than either routed approach.

The reason is specific to what dominates this corpus: scans account for about 94% of the bytes, and they are 150dpi. /ebook targets 150dpi, so it never downsamples them — Ghostscript only downsamples when the source exceeds roughly 1.5× the target. The loop therefore spends most of its runtime rewriting the files that hold nearly all the weight, without touching a single pixel of them.

It is worth being precise here, because the obvious explanation is wrong: /ebook does not bloat the text files. On my corpus it shrank them — by 91.5% on the ones with unsubsetted embedded fonts, because font subsetting is exactly what it does well. The naive loop’s failure is not that it damages text; it is that it does nothing for the scans that make up the bulk.

Why routing by images alone isn’t enough

Strategy ② sends anything containing images to Ghostscript and everything else to qpdf. That single rule captures most of the win — 26% to 82% — and it is the fastest of the three, because on my files qpdf finished in 12–22 ms against Ghostscript’s 180–400 ms, and most files take the qpdf path.

That speed gap is not a constant, though. --recompress-flate makes qpdf re-compress every flate stream at level 9, so on a PDF whose image data is flate-encoded rather than JPEG, qpdf has real work to do and its advantage narrows sharply. Drop the flag if throughput matters more than the last percent.

But it leaves 4 percentage points behind, and the reason is fonts. Ghostscript subsets embedded fonts; qpdf cannot. A text document that ships a full unsubsetted typeface has a large win available that has nothing to do with images — in my earlier tests, 73 KB down to 17 KB. Strategy ② routes those files to qpdf and collects a few percent instead of seventy.

Strategy ③ adds one check: if the file has no images but does carry embedded, unsubsetted fonts, send it to Ghostscript anyway.

The trade-off nobody mentions

Strategy ③ compresses best but takes twice as long as ②, because those 15 font-heavy files move from a ~20 ms qpdf call to a ~200 ms Ghostscript call. That is the honest cost:

If you care about Use Result here
Smallest output Strategy ③ −85.8%, 9.90 s
Best size per second Strategy ② −81.8%, 5.03 s
Neither Strategy ① −26.2%, 11.99 s

Extrapolated to 10,000 files of this mix, strategy ① costs roughly 50 minutes and returns a quarter of the savings; strategy ② about 21 minutes; strategy ③ about 41 minutes for the extra four points. Those figures assume the same composition — change the mix and they move.

Strategy ① is not universally wrong, and it is worth saying so plainly. On a folder of nothing but 300dpi scans it becomes competitive and can win: every file needs Ghostscript anyway, so the routing probes in ② and ③ are pure overhead, and /ebook leaves those scans at 150dpi where /screen drops them to 72dpi — often unreadable for archival documents. It also needs only one tool installed, which matters in locked-down or container environments.

The routing strategies win when the folder is mixed. That is the claim this benchmark supports; a scan-only archive is a different problem with a different answer.

The script

Strategy ③, as a plain bash loop. It needs poppler-utils for pdfimages and pdffonts:

#!/usr/bin/env bash
# Note: no `set -e`. qpdf exits 3 on warnings — a single damaged
# cross-reference table would otherwise kill the whole batch.
set -uo pipefail
mkdir -p out
processed=0; skipped=0

for f in *.pdf; do
  # Unreadable or password-protected? Skip it, don't abort the run.
  if ! pdfinfo "$f" >/dev/null 2>&1; then
    echo "SKIP (unreadable or encrypted): $f" >&2
    skipped=$((skipped+1)); continue
  fi

  if [ "$(pdfimages -list "$f" 2>/dev/null | tail -n +3 | wc -l)" -gt 0 ]; then
    gs -sDEVICE=pdfwrite -dPDFSETTINGS=/screen -dNOPAUSE -dQUIET -dBATCH -sOutputFile="out/$f" "$f" 2>/dev/null
  elif pdffonts "$f" 2>/dev/null | tail -n +3 | awk '$(NF-4)=="yes" && $(NF-3)=="no" { found=1 } END { exit !found }'; then
    gs -sDEVICE=pdfwrite -dPDFSETTINGS=/ebook -dNOPAUSE -dQUIET -dBATCH -sOutputFile="out/$f" "$f" 2>/dev/null
  else
    # qpdf: 0 = ok, 3 = warnings (output is still written), 2 = error
    qpdf --object-streams=generate --recompress-flate --compression-level=9 "$f" "out/$f" 2>/dev/null
    rc=$?
    if [ "$rc" -ne 0 ] && [ "$rc" -ne 3 ]; then
      echo "FAILED (qpdf $rc): $f" >&2
      skipped=$((skipped+1)); continue
    fi
  fi

  if [ -s "out/$f" ]; then processed=$((processed+1))
  else echo "FAILED (no output): $f" >&2; skipped=$((skipped+1)); fi
done

echo "processed=$processed skipped=$skipped"

One detail in that awk line is worth pausing on, because the obvious version is wrong. Writing { if (...) exit 0 } END { exit 1 } looks correct and always reports “no match” — in awk, exit still runs the END block, so the exit 1 overwrites your success code. Setting a flag and testing it in END is the version that works. I shipped the broken one first and only caught it because the font-heavy test file came out 4% smaller instead of 76%.

The set -e omission is deliberate, and it is the difference between a script that works on your test folder and one that works on a real archive. qpdf returns exit code 3 for warnings — a damaged cross-reference table produces output perfectly well but still reports 3, and damaged xrefs are common in scanned archives. With set -e, the first such file kills the loop. I tested it: nine files with one damaged PDF among them, and the run stopped after two, leaving seven unprocessed and an out/ directory that looks like a partial success. Encrypted files do the same thing via exit 2.

Two notes on the preset choice. Scans get /screen rather than /ebook because /ebook targets 150dpi and silently does nothing to a 150dpi source — Ghostscript only downsamples when the source exceeds about 1.5× the target. Font-only files get /ebook because there are no images to protect and the win comes from subsetting.

If your scans are consistently 300dpi or higher, /ebook is the better choice for them — it preserves more quality and still cut my 300dpi test file by 84.7%.

Before running it on anything you care about

  • Never write over your originals. The script writes to out/ for a reason. Compression is lossy for images and there is no undo.
  • Spot-check the output. Downsampling to 72dpi is aggressive; open a few results before deleting anything.
  • Compare sizes and keep the winner. On an unfamiliar corpus, running both tools on a sample of twenty files and comparing totals costs a few minutes and tells you more than any rule of thumb.
  • Watch for files that grow. If an output is larger than its input, the routing sent it to the wrong tool — that is the signal to check what the file actually contains.

Where the numbers come from

The per-file behavior behind these results is measured in two companion pieces: qpdf vs Ghostscript, head to head covers which tool wins on which file type and why, and Ghostscript /ebook not reducing PDF size explains the 1.5× downsampling threshold that makes preset choice matter so much.


Test setup. Measured 26 July 2026 · Ghostscript 10.05.1 · qpdf 12.2.0 · poppler-utils · Linux, single core. Corpus: 40 files, 19.18 MB, composed as described above from synthetic PostScript-generated documents and grayscale scan-style PDFs. Times are wall-clock for serial execution on one core; a multi-core machine running these jobs in parallel will be faster across the board, but the ranking between strategies is set by which tool each file is sent to, not by concurrency. Sizes are actual output totals.

References: qpdf manual — exit status (the documented meaning of exit code 3, which is why set -e breaks this loop) · Ghostscript High Level Devices (PDFSETTINGS presets and ImageDownsampleThreshold)