arXiv:2509.21057cs.CRcs.CL2025-09被引 15

提出PMark方法,实现无失真且抗改写攻击的语义级水印。

PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraints

  • 基于代理函数构建理论框架,动态估计句子中位数并约束多通道水印信号。
  • 在保持输出质量不变的前提下,对改写攻击的鲁棒性显著优于现有方法。
  • 适合需要高可靠性文本溯源的场景,如AI生成内容检测。

语义级水印(SWM)通过将句子作为基本单元,提升大语言模型生成文本对修改和改写攻击的鲁棒性。然而,现有方法缺乏强理论保障,且基于拒绝采样的生成方式常引入显著分布偏差。本文提出一种新的理论框架,引入代理函数(PF)——将句子映射为标量值的函数。基于此,我们提出PMark方法:通过采样动态估计下一句子的PF中位数,并施加多个PF约束(称为通道),强化水印证据。该方法具备坚实的理论保证,实现了无失真特性,并提升了对抗改写攻击的鲁棒性。此外,我们还设计了一种经实证优化的版本,无需动态中位数估计,进一步提高采样效率。实验表明,PMark在文本质量和鲁棒性上均持续优于现有SWM基线,为机器生成文本检测提供了更优范式。代码将发布于[GitHub链接]。

原文摘要 · Abstract (English)

Semantic-level watermarking (SWM) for large language models (LLMs) enhances watermarking robustness against text modifications and paraphrasing attacks by treating the sentence as the fundamental unit. However, existing methods still lack strong theoretical guarantees of robustness, and reject-sampling-based generation often introduces significant distribution distortions compared with unwatermarked outputs. In this work, we introduce a new theoretical framework on SWM through the concept of proxy functions (PFs) $\unicode{x2013}$ functions that map sentences to scalar values. Building on this framework, we propose PMark, a simple yet powerful SWM method that estimates the PF median for the next sentence dynamically through sampling while enforcing multiple PF constraints (which we call channels) to strengthen watermark evidence. Equipped with solid theoretical guarantees, PMark achieves the desired distortion-free property and improves the robustness against paraphrasing-style attacks. We also provide an empirically optimized version that further removes the requirement for dynamical median estimation for better sampling efficiency. Experimental results show that PMark consistently outperforms existing SWM baselines in both text quality and robustness, offering a more effective paradigm for detecting machine-generated text. Our code will be released at [this URL](https://github.com/PMark-repo/PMark).

语义水印大模型安全生成质量抗改写

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。