arXiv:2602.11793cs.CRcs.CL2026-02被引 2

弱水印组合比强水印更有效,能避免信号衰减。

More Haste, Less Speed: Weaker Single-Layer Watermark Improves Distortion-Free Watermark Ensembles

  • 用较弱单层水印保留词分布熵,利于多层组合。
  • 实验表明检测率和鲁棒性均优于强水印基线。
  • 适合关注生成内容溯源的模型开发者。

水印技术已成为检测和归属大语言模型生成内容的关键手段。尽管近期研究采用水印集成提升鲁棒性,但主流方法仍追求每层水印强度最大化。本文发现这种‘越强越好’策略存在根本缺陷:强水印显著降低词元分布熵,反而削弱后续层的水印效果。理论与实证表明,可检测性受熵值限制,水印集成导致熵与期望绿列表比例随层递减。为此,我们提出新框架,采用较弱的单层水印以维持多层集成所需的熵。实验验证该反直觉策略有效缓解信号衰减,在检测率与鲁棒性上持续优于强基线。

原文摘要 · Abstract (English)

Watermarking has emerged as a crucial technique for detecting and attributing content generated by large language models. While recent advancements have utilized watermark ensembles to enhance robustness, prevailing methods typically prioritize maximizing the strength of the watermark at every individual layer. In this work, we identify a critical limitation in this "stronger-is-better" approach: strong watermarks significantly reduce the entropy of the token distribution, which paradoxically weakens the effectiveness of watermarking in subsequent layers. We theoretically and empirically show that detectability is bounded by entropy and that watermark ensembles induce a monotonic decrease in both entropy and the expected green-list ratio across layers. To address this inherent trade-off, we propose a general framework that utilizes weaker single-layer watermarks to preserve the entropy required for effective multi-layer ensembling. Empirical evaluations demonstrate that this counter-intuitive strategy mitigates signal decay and consistently outperforms strong baselines in both detectability and robustness.

水印LLM熵控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。