元数据 · Metadata
PDF 元数据保留:合并时该注意什么 / PDF Metadata Preservation: What Survives a Merge
合并 PDF 时,元数据不会自动跟着走。文档信息字典在一个文件里只有一份,两份合成一份之后,标题、作者、主题、关键词只能留下其中一套值。动手之前先把每份文件的元数据抄下来,合并之后再回填你要的那一份。
阅读约 7 分钟 · 7 min read
一句话结论
PDF 元数据保留这件事,指望合并工具自动汇总是不行的。描述整个文档的字段在一个文件里只有一份,合并后只能留下其中一套值,多数工具取第一个源文件的那套。能完整保住的是绑在页面上的信息,页面尺寸、旋转角度这些跟着页面走。做法很简单:合并前把每份文件的标题、作者、关键词抄下来,合并后打开成品的属性面板逐项回填。
A merge does not carry metadata across on its own. The fields that describe a whole document exist once per file, so after two files become one only one set of values can stay, and most tools keep the first source. What does carry over intact is page level information, since media box and rotation travel with the page. Copy the title, author and keywords down before you merge, then write them back into the finished file.
元数据有哪几层 / Where metadata actually lives
PDF 的元数据不是一整块。它散在几个地方,合并时各自的下场也不一样,先分清位置再谈保留。
- 文档信息字典:标题、作者、主题、关键词、创建程序、创建与修改时间。一个文件只有一份
- XMP 包:XML 形式的一套扩展字段,版权状态、来源、标签常放在这里
- 页面级信息:页面尺寸、旋转角度、缩略图设置,跟着页面走
- 结构信息:书签、内部链接、表单字段,属于文档树的一部分
- 嵌入附件与输出意图:跟着具体对象走,合并时通常一起搬过去
PDF metadata is not a single block. It sits in several places and each one behaves differently during a merge, so it helps to know where a field lives before arguing about whether it should survive. The document information dictionary is the one people notice, because that is what the properties panel shows. XMP is the one that quietly disagrees with it.
哪些留下、哪些被覆盖 / What survives and what gets overwritten
下面用两份各带完整元数据的文件举例。合并前的取值、合并后的常见结果、以及该不该处理,都在表里。
| 字段 | 合并前 | 合并后常见结果 | 怎么处理 |
|---|---|---|---|
| 标题 / Title | 两份不同 | 只留一份 | 合并后手动回填 |
| 作者 / Author | 两份不同 | 被覆盖 | 可写成两个人 |
| 关键词 / Keywords | 两份不同 | 被覆盖 | 合并去重后回填 |
| 创建时间 / CreationDate | 各不相同 | 取其一 | 以最早那份为准 |
| 修改时间 / ModDate | 各不相同 | 变成合并当天 | 正常,不用改 |
| 创建程序 / Producer | 可能不同 | 变成合并工具 | 正常,不用改 |
| 页面尺寸与旋转 | 各不相同 | 逐页保留 | 绑在页面上,不用管 |
| XMP 扩展字段 | 两份不同 | 留一份或整包丢弃 | 改完信息字典后重新生成 |
表里最该盯的是前三项加 XMP。修改时间和创建程序变化属于正常记录,不用去改;标题、作者、关键词变了才是真的丢了东西。
The rows that matter are the first three plus XMP. A changed modification date or producer is just the record of what happened, and there is nothing to restore. A changed title, author or keyword list means information is genuinely gone from the finished file.
三个最容易丢的场合 / Three ways metadata goes missing
- 指望元数据自动汇总。不会。合并工具只挑一份写进成品,通常取第一个源文件,剩下的那套直接丢掉。
- 先压缩再合并。不少压缩工具会顺手清掉元数据字段。顺序换成先合并再压缩,压缩只做一次。
- 只改信息字典不改 XMP。两套并存且不同步,只改一边会让不同工具读到不同结果,检索系统往往读到旧的那一套。
Expecting metadata to combine automatically is the common one, and it never works: the merge tool picks one set and the rest is dropped. Compressing before merging is the sneaky one, since plenty of compression tools strip fields as a side effect, so merge first and compress once. Editing the info dictionary while leaving XMP alone leaves the two disagreeing, and search systems tend to read whichever one you did not touch.
合并前后怎么回填 / A workflow that keeps the fields you care about
把抄和回填两步放进流程里,成本只有一两分钟,比事后从源文件里翻回原值省事得多。文件全程不出本机。
- 合并前逐份打开属性面板,把标题、作者、关键词抄下来
- 判断哪一份的元数据描述的是成品,作为回填的底稿
- 按最终阅读顺序排好源文件,一次合并,避免事后返工
- 合并后打开成品属性面板,逐项对照抄下来的值
- 回填标题、作者、关键词,作者可以两份合一,关键词去重后一次填入
- 创建时间取最早那份,修改时间保持为合并当天
- 检查 XMP 与信息字典是否一致,不一致就重新生成一份
- 成品与母版分开存,母版不动
书签和表单字段的保留方式见 /blog/pdf-bookmark-merge-keep-outline 与 /blog/pdf-form-merge-keep-fields;压缩与合并的先后顺序在 /blog/compressed-pdf-merge-quality-loss 里有说明。
在 pdfmergenext.shop 合并 PDF,合并前先抄下元数据 →更多 PDF 元数据保留的做法见 /blog。
常见问题 / FAQ
合并后标题和作者怎么只剩一个?
因为文档信息字典在一个文件里只有一份。合并工具会挑一份写进成品,通常是第一个源文件的那一份。要保住两份的内容,就在合并后手动回填,作者可以写成两个人,关键词合并去重后一次填入。
修改时间变了要紧吗?
不要紧。修改时间记录的是成品最后一次被写盘的时间,它就该是合并那天。真正需要回填的是标题、作者、关键词这类描述性字段,它们描述的是文件内容而不是操作时间。
XMP 和文档信息字典要同时改吗?
要。两套并存且不会自动同步,只改一边会让不同工具读到不同结果。改完一边就重新生成另一边,或者直接用能同时写两边的工具处理一次。
Why does the merged file only keep one title and one author?
The document information dictionary holds one set of values per file. A merge tool picks one and writes it into the output, usually the first source. Write the values back by hand if you need both, combining the author names and de-duplicating the keywords.
Does it matter that the modification date changed?
No. That field records the last write to the finished file, and the merge day is the correct value for it. Title, author and keywords are the fields worth restoring, since those describe the content rather than the operation.
Do I have to update XMP and the info dictionary together?
Yes. Both exist side by side and neither syncs automatically, so editing only one leaves different tools reading different values. Regenerate the other side after each edit, or use a tool that writes both in one pass.