PDF 合并

压缩质量 · Compression Quality

PDF 压缩后合并:质量损失有多大 / Compressed PDF Merge: How Much Quality Do You Lose

做 PDF 压缩后合并时,大家担心的是再压缩一次会不会把图压糊。答案是合并这一步基本不动画质,问题出在同一份图像走的两次有损压缩。这篇把顺序讲清楚。

阅读约 7 分钟 · 7 min read

一句话结论

压缩过的 PDF 可以直接合并,合并本身不会带来 PDF 压缩后合并的额外质量损失。画质保不住是因为同一张图经历了两次有损压缩:源文件压过一次,合并后再压一次。把顺序改成先合并、最后统一压一次,损失就只剩一轮。

A merge copies page objects into one container and leaves image streams as they are, so quality does not drop on the way in. What hurts is running lossy compression twice. Merge first, then compress the combined file once, and you spend a single pass instead of several.

合并到底动了画质没有 / What a merge does to image quality

PDF 里的照片大多是一段已经编码好的字节流,JPEG 或 JPEG2000 居多。合并工具读的是这串字节,把它连同页面描述一起搬进新文件,中间没有解码再编码的环节。既然没有重新编码,也就不会因为合并而丢像素。

  • 页面尺寸、旋转角度、书签层级这类结构信息按原样复制,不参与重新计算
  • 字体子集保持完整,文字继续可选可搜
  • 输出体积约等于两份源文件之和,多一点的部分是交叉引用表与页面树
  • 只有「打印成 PDF」「导出为图片」这类路径才会重新生成像素,画质这时候才开始掉

Most photos in a PDF are already encoded byte streams, usually JPEG or JPEG2000. A merge reads those bytes, copies them with the page description into a new file, and never decodes them, which is why no pixels are lost. Page geometry, rotation, outline nesting and font subsets all travel across intact. The size of the result is roughly the two sources added together, plus a little for the page tree. Reprinting or exporting to an image is what regenerates pixels, and that is where sharpness starts to go.

先压缩还是先合并:顺序决定最终画质 / Compress before or after the merge

做法有损压缩次数画质结果建议
每个源文件先压,再合并两次画质损失最明显不推荐
直接合并,之后不再压零次画质完全保留源文件已压缩时的首选
先合并,最后统一压一次一次可接受,体积能降推荐路径
合并后压,之后又改并重压多次逐次劣化务必避免
  • 源文件已经压过一轮,且体积在可接受范围内:直接合并,不再压缩
  • 合并后体积超标(邮件附件、系统上传限制):先合并,最后统一压一次
  • 源文件清晰度参差不齐:先把扫描页统一到同一分辨率,再合并
  • 文件还要反复修改:留着合并前的原件,每次从原件重新出稿,不要在成品上反复压

When the sources already went through compression once and the size is acceptable, merge them and stop. If the combined file is over a limit, compress once at the end. Match scan resolutions before merging so the pages look even, and keep the pre-merge originals around, because re-compressing an already compressed result is what stacks the damage.

二次压缩为什么会糊:误差第一次 / Why the second pass is the expensive one

有损压缩的做法是按小块丢弃人眼不太敏感的高频细节。第一轮结束时图像已经不是原始数据,第二轮在这个已经有偏差的版本上再判断哪些细节可以丢,丢出来的东西和原始画面的偏离就被放大了。文字边缘、细线、笔划拐角最先显出来。

  • 第一遍写字边变软,第二遍开始出现毛刺,第三遍直接结块
  • 渐变区域从平滑变成一段一段的色带,天空和灯光背景最明显
  • 体积收益递减:第二轮压缩通常只能再省几个百分点,画质的代价远大于这点体积
  • 扫描件尤其吃亏,字迹在第二遍之后开始发灰,识别率也跟着降

Lossy coding discards high-frequency blocks that are easy to miss. After the first pass the file no longer holds the original data, so the second pass decides what to throw away from an already degraded copy and pushes the result further from the source. Edges soften first, then rot into visible fringing, and gradients break into bands. The size savings shrink with each pass while the visual cost grows, which is why a second round is rarely worth it.

哪些内容经得起压缩,哪些经不起 / What survives compression

内容类型压缩后的表现怎么处理
矢量文字层 / Vector text几乎看不出变化放心压,压缩对矢量不产生块效应
照片渐变(天空、灯光)容易出现色带降分辨率优于提高压缩比
图表细线 / Thin strokes线条断裂、边缘发毛保留原始 PDF,别用截图代替
扫描页面 / Scanned pages每压一次字迹更糊先降到 200 DPI,或改用二值化
已压过的 JPEG 图画质掉得最快,体积几乎不减直接合并,跳过第二次压缩
透明图层 / Transparency可能被扁平化压缩前确认已是最终版

矢量文字层基本免疫,因为压缩动的是 raster 图像,轮廓是算出来的。真正需要盯住的是照片、扫描页和细线图,这三类占据了大部分「压缩后看不清」的投诉。

Vector text is largely immune, since compression targets raster images while outlines are drawn from math. Photos, scanned pages and thin-line charts account for most complaints about unreadable output, and they deserve a look before you hit the button.

一次成型的本地流程 / A local workflow that compresses once

要先拼一版草稿给别人看,就先把原分辨率的文件合并出来。确认内容定稿之后,再对这个成品做唯一一次压缩。文件全程不出本机,也不用在多个在线工具之间来回上传。

  1. 检查每份源文件的分辨率,把扫描页统一到同一数值
  2. 按最终阅读顺序排列文件,一次排好,避免后续调整后重新压缩
  3. 在本地完成合并,先输出全分辨率版本存好
  4. 用阅读器翻一遍,确认页面顺序、矢量文字和图片清晰度没问题
  5. 内容定稿后,只对这版成品执行一次压缩
  6. 存压缩版用于发送,全分辨率版留着当作下一次修改的母版
  7. 压缩结果与母版各留一份单独命名的副本,别覆盖

要看本地压缩的取舍,可以看 /blog/compress-pdf-local-no-upload 里的说明;涉及几百 MB 的材料,先读 /blog/large-pdf-merge-500mb 会省不少时间。

在 pdfmergenext.shop 合并 PDF,压缩只做一次 →

更多 PDF 压缩后合并的做法见 /blog。

常见问题

压缩过的 PDF 合并之后还需要再压一次吗?

多数情况不需要。合并只是把已有页面对象搬到一个新容器,体积基本等于两份源文件相加。真要压,就在合并后压这一次,不要每个源文件各压一遍再合。

为什么二次压缩之后文字边缘发毛?

有损编码按块丢弃高频细节,字边和线条正好是高频部分。第一遍已经丢过一次,第二遍在丢过的基础上再丢,边缘就会出现毛刺和锯齿。

合并后文件太大,怎么压最不容易掉画质?

优先降 raster 图像的分辨率而不是提高压缩比。把 300 DPI 的扫描页降到 200 DPI,文字通常还能看清,画质损失比再压一轮 JPEG 小得多。

PDF 里的透明图层会被压缩弄没吗?

有可能。部分压缩路径会把透明元素扁平化成一张位图,背景就没法保持透明了。合成之前先把这类文件看一眼,确认拼的是最终稿。

Does merging PDFs reduce image quality?

Not on its own. A merge copies already encoded image streams into a new document and does not re-encode them. Visible softening comes from running lossy compression a second time, either before or after the merge.

Should I compress each file first, or merge first and then compress?

Merge first, then compress once. Compressing every source separately and merging afterwards means every image goes through two lossy passes, which costs more detail than one pass on the combined file.

相关阅读 / Related