arXiv:2505.05190cs.LGcs.AI2025-05ICML被引 21

攻击者利用文本水印的高熵设计漏洞,实现低成本高效擦除。

Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks

  • 通过计算词元自信息,精准定位可篡改的水印标记。
  • 对7种主流水印方法攻击成功率接近100%,每百万词仅需0.88美元。
  • 无需模型或算法信息,可迁移至任意大模型甚至手机端使用。

文本水印通过控制大语言模型(LLM)采样过程,在文本中嵌入统计信号,使检测器能验证输出是否来自指定模型。当前水印算法将水印嵌入高熵词元以保证文本质量,但本文揭示该设计存在严重漏洞。我们提出通用高效的改写攻击——自信息重写攻击(SIRA),通过计算每个词元的自信息,识别潜在的模式词元并实施精准攻击。实验表明,SIRA在七种近期水印方法上均达到近100%成功率,成本仅为每百万词0.88美元。该方法无需访问水印算法或水印模型,可无缝迁移至任意大语言模型,包括移动端模型。研究凸显了现有水印技术脆弱性,亟需更鲁棒的解决方案。

原文摘要 · Abstract (English)

Text watermarking aims to subtly embed statistical signals into text by controlling the Large Language Model (LLM)'s sampling process, enabling watermark detectors to verify that the output was generated by the specified model. The robustness of these watermarking algorithms has become a key factor in evaluating their effectiveness. Current text watermarking algorithms embed watermarks in high-entropy tokens to ensure text quality. In this paper, we reveal that this seemingly benign design can be exploited by attackers, posing a significant risk to the robustness of the watermark. We introduce a generic efficient paraphrasing attack, the Self-Information Rewrite Attack (SIRA), which leverages the vulnerability by calculating the self-information of each token to identify potential pattern tokens and perform targeted attack. Our work exposes a widely prevalent vulnerability in current watermarking algorithms. The experimental results show SIRA achieves nearly 100% attack success rates on seven recent watermarking methods with only 0.88 USD per million tokens cost. Our approach does not require any access to the watermark algorithms or the watermarked LLM and can seamlessly transfer to any LLM as the attack model, even mobile-level models. Our findings highlight the urgent need for more robust watermarking.

文本水印攻击方法大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。