arXiv:2502.11268cs.CL2025-02ACL被引 19

提出多通道无偏水印技术,提升大模型文本水印检测率10%以上。

Improved Unbiased Watermark for Large Language Models

  • 将词汇表分段,按密钥选择性增强特定段落的词概率
  • 相比现有技术,检测率提升超10%,且不破坏生成文本质量
  • 适合需要可验证生成来源的AI内容平台使用

随着人工智能在文本生成方面超越人类能力,验证AI生成内容的来源变得至关重要。无偏水印通过在语言模型生成的文本中嵌入统计信号,无需影响文本质量即可实现溯源。本文提出MCmark,一种基于多通道的无偏水印方法。MCmark将模型词汇表划分为若干段,根据水印密钥提升选定段内词的概率。实验表明,MCmark不仅保持了语言模型原有的分布特性,还在检测率和鲁棒性上显著优于现有无偏水印技术。在多个主流大语言模型上的测试显示,相较当前最先进的无偏水印,其检测率提升超过10%。这一进展凸显了MCmark在实际应用中的潜力。

原文摘要 · Abstract (English)

As artificial intelligence surpasses human capabilities in text generation, the necessity to authenticate the origins of AI-generated content has become paramount. Unbiased watermarks offer a powerful solution by embedding statistical signals into language model-generated text without distorting the quality. In this paper, we introduce MCmark, a family of unbiased, Multi-Channel-based watermarks. MCmark works by partitioning the model's vocabulary into segments and promoting token probabilities within a selected segment based on a watermark key. We demonstrate that MCmark not only preserves the original distribution of the language model but also offers significant improvements in detectability and robustness over existing unbiased watermarks. Our experiments with widely-used language models demonstrate an improvement in detectability of over 10% using MCmark, compared to existing state-of-the-art unbiased watermarks. This advancement underscores MCmark's potential in enhancing the practical application of watermarking in AI-generated texts.

水印技术大模型文本溯源安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。