arXiv:2411.01222cs.CL2024-11NAACL被引 8

提出黑盒攻击方法B⁴,有效清除大模型文本水印

$B^4$: A Black-Box Scrubbing Attack on LLM Watermarks

  • 将水印清除建模为约束优化问题,用两个分布描述目标
  • 在12种设置下优于现有基线方法,清除效果显著
  • 无需了解水印细节,适合真实场景下的对抗攻击研究

水印技术已成为检测大模型生成内容的主流方法,通过嵌入难以察觉的模式实现。尽管性能优异,其对抗攻击的鲁棒性仍缺乏充分研究。以往工作多假设灰盒攻击环境,需知晓水印类型甚至超参数,这在实际中难以满足。针对更贴近现实的黑盒威胁模型,本文提出B⁴——一种无需先验知识的黑盒水印擦除攻击。通过构建水印分布与保真度分布来形式化攻击目标,将水印清除转化为约束优化问题,并采用代理分布近似求解。在12种不同设置下的实验表明,B⁴在清除效果上显著优于其他基线方法。

原文摘要 · Abstract (English)

Watermarking has emerged as a prominent technique for LLM-generated content detection by embedding imperceptible patterns. Despite supreme performance, its robustness against adversarial attacks remains underexplored. Previous work typically considers a grey-box attack setting, where the specific type of watermark is already known. Some even necessitates knowledge about hyperparameters of the watermarking method. Such prerequisites are unattainable in real-world scenarios. Targeting at a more realistic black-box threat model with fewer assumptions, we here propose $B^4$, a black-box scrubbing attack on watermarks. Specifically, we formulate the watermark scrubbing attack as a constrained optimization problem by capturing its objectives with two distributions, a Watermark Distribution and a Fidelity Distribution. This optimization problem can be approximately solved using two proxy distributions. Experimental results across 12 different settings demonstrate the superior performance of $B^4$ compared with other baselines.

大模型安全水印攻击黑盒攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。