用最大耦合技术让大模型水印更隐蔽且难被消除
Debiasing Watermarks for Large Language Models via Maximal Coupling
- 将词汇分为绿红列表,通过均匀随机决定是否修正分布偏差
- 在保持文本质量的同时实现高检测率,抗修改能力强
- 适合需要防伪但又不破坏生成内容的场景
为区分人工与机器生成文本以维护数字通信的可信度,本文提出一种基于最大耦合的绿色/红色列表水印方法。该方法将词元集划分为‘绿色’和‘红色’列表,轻微提升绿色词元的生成概率。为纠正词元分布偏差,采用最大耦合策略,利用均匀硬币翻转决定是否应用偏差修正,结果作为伪随机水印信号嵌入。理论分析证明该方法具有无偏性和强鲁棒性。实验表明,相比先前技术,在保持文本质量的同时实现更高可检测性,并对旨在提升文本质量的针对性修改表现出强抵抗力。本研究为语言模型提供了一种有效检测且对文本质量影响极小的水印解决方案。
原文摘要 · Abstract (English)
Watermarking language models is essential for distinguishing between human and machine-generated text and thus maintaining the integrity and trustworthiness of digital communication. We present a novel green/red list watermarking approach that partitions the token set into ``green'' and ``red'' lists, subtly increasing the generation probability for green tokens. To correct token distribution bias, our method employs maximal coupling, using a uniform coin flip to decide whether to apply bias correction, with the result embedded as a pseudorandom watermark signal. Theoretical analysis confirms this approach's unbiased nature and robust detection capabilities. Experimental results show that it outperforms prior techniques by preserving text quality while maintaining high detectability, and it demonstrates resilience to targeted modifications aimed at improving text quality. This research provides a promising watermarking solution for language models, balancing effective detection with minimal impact on text quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。