arXiv:2602.06754cs.CRcs.AI2026-02被引 4

统一了大模型水印方法,揭示了质量、多样性与可检测性的权衡关系。

A Unified Framework for LLM Watermarks

  • 从约束优化视角统一现有水印算法,揭示其内在机制。
  • 不同约束下设计的水印方案均显著提升可检测性。
  • 适用于需要定制化水印的场景,如注重生成质量时。

大模型水印可通过在生成文本中插入可检测信号实现内容溯源。现有水印方法设计各异,多采用自下而上的方式,缺乏通用且有理论基础的框架。本文表明,多数现有水印方案均可由一个统一的约束优化问题推导得出。该框架不仅整合了已有方法,还显式揭示了每种方法所优化的约束条件。特别地,它揭示了一个被忽视的质量-多样性-可检测性三者权衡关系。同时,该框架为定制新型水印方案提供了理论指导,例如可直接以困惑度(perplexity)作为质量代理指标,并推导出针对该约束最优的新方案。实验验证表明,基于特定约束推导的水印方案,在对应约束下始终具备最强的可检测性。

原文摘要 · Abstract (English)

LLM watermarks allow tracing AI-generated texts by inserting a detectable signal into their generated content. Recent works have proposed a wide range of watermarking algorithms, each with distinct designs, usually built using a bottom-up approach. Crucially, there is no general and principled formulation for LLM watermarking. In this work, we show that most existing and widely used watermarking schemes can in fact be derived from a principled constrained optimization problem. Our formulation unifies existing watermarking methods and explicitly reveals the constraints that each method optimizes. In particular, it highlights an understudied quality-diversity-power trade-off. At the same time, our framework also provides a principled approach for designing novel watermarking schemes tailored to specific requirements. For instance, it allows us to directly use perplexity as a proxy for quality, and derive new schemes that are optimal with respect to this constraint. Our experimental evaluation validates our framework: watermarking schemes derived from a given constraint consistently maximize detection power with respect to that constraint.

大模型水印生成安全约束优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。