arXiv:2507.06274cs.CRcs.AI2025-07NeurIPS被引 6

提出新水印机制,同时增强抗删除和伪造攻击能力。

Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks

  • 用多重独立检测令牌实现水印冗余,突破传统权衡。
  • 在多个数据集上,伪造攻击鲁棒性提升82%~92%,删除攻击提升6.4%~24.6%。
  • 适合需要高安全性的大模型内容溯源场景。

水印是防范大语言模型滥用的有力手段,但仍易受删除与伪造攻击。其脆弱性源于水印窗口大小的固有权衡:窗口越小,越抗删除,却越易被低成本统计方法逆向破解。本文提出等效纹理密钥机制,使窗口内多个标记可独立支持检测,利用冗余实现突破。基于此设计了子词汇分解等效纹理密钥(SEEK)水印方案,在不降低伪造攻击防御力的前提下,显著提升对删除攻击的鲁棒性。实验表明,相较已有方法,SEEK在不同数据集下,伪造攻击鲁棒性提升88.2%、92.3%、82.0%,删除攻击鲁棒性提升10.2%、6.4%、24.6%。

原文摘要 · Abstract (English)

Watermarking is a promising defense against the misuse of large language models (LLMs), yet it remains vulnerable to scrubbing and spoofing attacks. This vulnerability stems from an inherent trade-off governed by watermark window size: smaller windows resist scrubbing better but are easier to reverse-engineer, enabling low-cost statistics-based spoofing attacks. This work breaks this trade-off by introducing a novel mechanism, equivalent texture keys, where multiple tokens within a watermark window can independently support the detection. Based on the redundancy, we propose a novel watermark scheme with Sub-vocabulary decomposed Equivalent tExture Key (SEEK). It achieves a Pareto improvement, increasing the resilience against scrubbing attacks without compromising robustness to spoofing. Experiments demonstrate SEEK's superiority over prior method, yielding spoofing robustness gains of +88.2%/+92.3%/+82.0% and scrubbing robustness gains of +10.2%/+6.4%/+24.6% across diverse dataset settings.

水印技术LLM安全抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。