字符级扰动可高效破坏大模型水印,暴露现有方案漏洞。
Character-Level Perturbations Disrupt LLM Watermarks
- 通过字符级错误扰乱分词,一次修改影响多个词元
- 在受限查询条件下,基因算法优化攻击成功率超传统方法
- 揭示固定防御易被绕过,适合安全与对抗研究者参考
大型语言模型(LLM)水印技术通过嵌入可检测信号实现版权保护、滥用防范和内容识别。现有研究多基于水印移除攻击评估鲁棒性,但这些方法往往效率低下,导致误判认为有效移除需大幅扰动或强敌手。为此,我们首先形式化了LLM水印系统模型,并刻画了两种在有限水印检测器访问权限下的现实威胁模型。分析发现,字符级扰动(如拼写错误、字符互换、删除、形似字符)可通过破坏分词过程同时影响多个词元,其攻击范围远超预期。实验表明,在最严格的威胁模型下,字符级扰动在水印移除上显著更有效。我们进一步提出基于遗传算法(GA)的引导式移除攻击,利用参考检测器进行优化,在仅限黑盒查询的实用场景中表现优异。结果验证了字符级扰动的优势及GA的有效性。此外,我们指出防御存在对抗困境:任何固定防御都可能被特定扰动策略绕过。据此,提出自适应复合字符级攻击,实验证明其可有效突破现有防御。研究揭示了现有水印机制的重大漏洞,凸显亟需构建新型鲁棒水印方案。
原文摘要 · Abstract (English)
Large Language Model (LLM) watermarking embeds detectable signals into generated text for copyright protection, misuse prevention, and content detection. While prior studies evaluate robustness using watermark removal attacks, these methods are often suboptimal, creating the misconception that effective removal requires large perturbations or powerful adversaries. To bridge the gap, we first formalize the system model for LLM watermark, and characterize two realistic threat models constrained on limited access to the watermark detector. We then analyze how different types of perturbation vary in their attack range, i.e., the number of tokens they can affect with a single edit. We observe that character-level perturbations (e.g., typos, swaps, deletions, homoglyphs) can influence multiple tokens simultaneously by disrupting the tokenization process. We demonstrate that character-level perturbations are significantly more effective for watermark removal under the most restrictive threat model. We further propose guided removal attacks based on the Genetic Algorithm (GA) that uses a reference detector for optimization. Under a practical threat model with limited black-box queries to the watermark detector, our method demonstrates strong removal performance. Experiments confirm the superiority of character-level perturbations and the effectiveness of the GA in removing watermarks under realistic constraints. Additionally, we argue there is an adversarial dilemma when considering potential defenses: any fixed defense can be bypassed by a suitable perturbation strategy. Motivated by this principle, we propose an adaptive compound character-level attack. Experimental results show that this approach can effectively defeat the defenses. Our findings highlight significant vulnerabilities in existing LLM watermark schemes and underline the urgency for the development of new robust mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。