arXiv:2411.05277cs.CRcs.CL2024-11EMNLP被引 29

仅用少量生成文本,就能破解主流水印方案的防改写能力。

Revisiting the Robustness of Watermarking to Paraphrasing Attacks

  • 通过黑盒获取少量生成文本,反推水印编码机制。
  • 对水印文本进行轻量级改写即可大幅降低检测率。
  • 适用于关注内容溯源安全的研究者与平台方。

随着语言模型生成内容在互联网上的泛滥,水印技术被视为验证文本是否由模型生成的可靠方法。近期许多水印方案通过微调语言模型输出概率来嵌入可检测信号。尽管部分方案声称具备抗改写能力,但其安全性并未充分考虑被逆向破解的可能。本文表明,仅需访问少量来自黑盒水印模型的生成文本,即可显著提升改写攻击的有效性,使水印几乎无法被检测,从而暴露现有方案的根本缺陷。

原文摘要 · Abstract (English)

Amidst rising concerns about the internet being proliferated with content generated from language models (LMs), watermarking is seen as a principled way to certify whether text was generated from a model. Many recent watermarking techniques slightly modify the output probabilities of LMs to embed a signal in the generated output that can later be detected. Since early proposals for text watermarking, questions about their robustness to paraphrasing have been prominently discussed. Lately, some techniques are deliberately designed and claimed to be robust to paraphrasing. However, such watermarking schemes do not adequately account for the ease with which they can be reverse-engineered. We show that with access to only a limited number of generations from a black-box watermarked model, we can drastically increase the effectiveness of paraphrasing attacks to evade watermark detection, thereby rendering the watermark ineffective.

水印对抗攻击语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。