arXiv:2607.13003cs.CRcs.IT2026-07

从信息论角度揭示生成模型水印的取证代价与机制。

Watermark Forensics for Generative Models: An Information-Theoretic Perspective

  • 用信息剖面函数量化每一步生成内容对秘密的泄露程度。
  • 多用户归属需Θ(log N/h)个词元,载荷提取需Θ(ℓ/h)个词元。
  • 首次给出精确的熵率定律,适用于大模型水印取证场景。

生成模型输出中的水印不仅可用于判断文本是否为机器生成,还可用于用户归属、隐藏信息提取及编辑后位置定位,构成一条取证阶梯。本文从信息论视角分析每级任务的样本长度代价。设秘密S为用户身份或隐藏数据,信息剖面ν(t)=I(S;X_t|X_{<t})记录第t个词元在已知前序词元条件下对S的贡献。其总质量决定归属与提取成本;分布形态决定定位精度;而检测仅依赖标记分布与原始分布的距离。现有两种水印模型——逐词微调或少数词元强标记——对应不同的剖面上限。主定理确立了阶梯的熵率维度:对无统计失真的方案,在熵率为h的平稳遍历源上,将文本归于N个用户需Θ(log N/h)个词元,精度达到(1+o(1))因子;这是首个通过精确对齐实现的多用户归属紧致熵率律。传统碰撞计数法会无限高估,唯有基于每个候选者自身实际惊讶度进行解码阈值判定,才能达到理论速率且几乎不误指无辜用户。反向证明使该律双向成立,载荷ℓ比特提取需Θ(ℓ/h)词元。两个真实差距存在:一个Θ(log N)词元的窗口内,文本可被证明为机生成但无法归属;另一个是足迹分辨率的不确定性原理。在GPT-2、Pythia-410M和Qwen2.5上的实验验证了预测常数。

原文摘要 · Abstract (English)

A watermark in a generative model's output is usually asked only whether a text is machine-made. The same mark can do more: attribute it to the user who produced it, extract a hidden payload, or localize the part that survives editing. These form a forensic ladder, and we ask what each rung costs in the sample length $n$. One object organizes the answers. Let $S$ be the secret the mark carries (a user's identity or payload), and let the information profile $ν(t)=I(S;X_t\mid X_{<t})$ record how much the $t$-th token reveals about $S$ given the earlier ones. Its total mass pays for attribution and extraction; how that mass is spread pays for localization; and detection alone is paid for not by information but by presence, the distance from the marked to the unmarked distribution. The literature's two quality models, a mark subtle on every token and one that stamps a few tokens loudly, are two incomparable ways of capping this profile. Our main theorem settles the ladder's entropy column. For statistically distortion-free schemes, attributing a text to one of $N$ users costs $Θ(\log N/h)$ tokens over every stationary-ergodic source of entropy rate $h$, sharp to a $(1+o(1))$ factor: to our knowledge the first tight entropy-rate law for multi-user attribution (via exact alignment). The natural collision-counting analysis overcharges without bound; only a decoder thresholding each candidate by its own realized surprisal attains the rate while almost never implicating an innocent user. A matching converse makes the law two-sided, and extraction of an $\ell$-bit payload costs $Θ(\ell/h)$. Two gaps are real, not modeling artifacts: a $Θ(\log N)$-token window in which a text is provably machine-made yet unattributable, and a footprint-resolution uncertainty principle. Experiments on GPT-2, Pythia-410M, and Qwen2.5 recover the predicted constants.

水印取证信息论生成模型隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。