arXiv:2502.18608cs.CRcs.LG2025-02

破解无损水印技术,实现对大模型生成内容的伪造攻击

Breaking Distortion-free Watermarks in Large Language Models

  • 通过自适应提示与排序算法逆向提取水印密钥
  • 在多个主流模型上成功实现大规模伪造文本生成
  • 挑战现有无损水印技术的鲁棒性理论假设

近年来,大语言模型水印技术成为防范人工智能生成内容的重要手段,具有广泛的应用前景。然而,当前水印方案面临专家级对手逆向破解的威胁。已有研究主要针对Kirchenbauer等人(2023)的分布扰动方法,而本文聚焦于更先进的无损水印技术(Kuditipudi等,2024),该技术通过隐藏的密钥序列保持原始词元分布不变。我们证明,即便在更复杂的水印机制下,仍可实现模型破坏并实施伪造攻击——生成大量可被归因于原始水印模型的(潜在有害)文本。具体而言,提出使用自适应提示与基于排序的算法,精确恢复水印密钥。在LLAMA-3.1-8B-Instruct、Mistral-7B-Instruct、Gemma-7b和OPT-125M上的实证结果,挑战了当前关于无损水印技术鲁棒性与可用性的理论主张。

原文摘要 · Abstract (English)

In recent years, LLM watermarking has emerged as an attractive safeguard against AI-generated content, with promising applications in many real-world domains. However, there are growing concerns that the current LLM watermarking schemes are vulnerable to expert adversaries wishing to reverse-engineer the watermarking mechanisms. Prior work in breaking or stealing LLM watermarks mainly focuses on the distribution-modifying algorithm of Kirchenbauer et al. (2023), which perturbs the logit vector before sampling. In this work, we focus on reverse-engineering the other prominent LLM watermarking scheme, distortion-free watermarking (Kuditipudi et al. 2024), which preserves the underlying token distribution by using a hidden watermarking key sequence. We demonstrate that, even under a more sophisticated watermarking scheme, it is possible to compromise the LLM and carry out a spoofing attack, i.e. generate a large number of (potentially harmful) texts that can be attributed to the original watermarked LLM. Specifically, we propose using adaptive prompting and a sorting-based algorithm to accurately recover the underlying secret key for watermarking the LLM. Our empirical findings on LLAMA-3.1-8B-Instruct, Mistral-7B-Instruct, Gemma-7b, and OPT-125M challenge the current theoretical claims on the robustness and usability of the distortion-free watermarking techniques.

大模型安全水印攻击逆向工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。