arXiv:2606.04486cs.CRcs.CL2026-06

提出一种针对扩散语言模型的全局水印方法,提升检测鲁棒性。

Global Sketch-Based Watermarking for Diffusion Language Models

  • 用全局向量草图控制文本生成,不依赖局部上下文。
  • 水印在不同顺序下仍可检测,且不产生简单词元偏移。
  • 适合需要抗篡改、跨序列验证的生成内容场景。

语言模型的水印技术在自回归生成中研究广泛,其多采用基于局部上下文的方法,通过前序词元调整下一个词元分布。而在扩散语言模型中,多个未确定位置的概率分布是联合采样的,使得整个序列的可加统计量在生成过程中可计算。本文提出一种针对掩码扩散语言模型的水印方法,通过控制文本的全局向量草图表示实现水印嵌入。相比依赖上下文的水印,该草图形式将检测与生成时的局部上下文解耦,形成顺序无关的统计特征,且水印规则不会表现为简单的词元偏置。我们分析了该方法的失真度、正确性和鲁棒性。

原文摘要 · Abstract (English)

Watermarking methods for language models have been studied extensively in the autoregressive setting, where tokens are generated sequentially. These works largely focus on local-context schemes that perturb the next token's distribution as a function of its preceding tokens. In diffusion language models, distributions over many unresolved positions are jointly sampled, allowing additive statistics of the entire sequence to be tractable during generation. We propose a watermark for masked diffusion language models that controls a global, vector-valued sketch representation of the text. Compared to context-dependent watermarking, the sketch formulation decouples detection from the local contexts seen during generation, resulting in an order-agnostic statistic and a watermarking rule which does not manifest as a simple token bias. We analyze the distortion, soundness, and robustness properties of the method.

扩散模型水印语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。