评估大模型水印在抗攻击与文本质量间的平衡
Signature vs. Substance: Evaluating the Balance of Adversarial Resistance and Linguistic Quality in Watermarking Large Language Models
- 对比改写与回译攻击测试水印鲁棒性
- 水印导致文风偏离且易被回译攻破
- 适合关注生成内容安全与真实性的研究者
为缓解大语言模型生成文本可能带来的危害,研究人员提出了水印技术,即在文本中嵌入可检测信号。通过水印可始终准确识别出由大模型生成的文本。然而,近期研究发现,此类技术常降低生成文本的质量,且对抗攻击可剥离水印信号,使文本逃避检测。这导致大模型开发者对水印技术的采纳产生犹豫。为此,我们评估了几种水印技术对对抗攻击的鲁棒性,比较了改写和回译(英→其他语言→英)攻击下的表现;并通过语言学指标评估其保持原始文本质量与写作风格的能力。结果表明,这些水印技术虽能保留语义,但明显偏离未加水印文本的写作风格,且对回译攻击尤为脆弱。
原文摘要 · Abstract (English)
To mitigate the potential harms of Large Language Models (LLMs)generated text, researchers have proposed watermarking, a process of embedding detectable signals within text. With watermarking, we can always accurately detect LLM-generated texts. However, recent findings suggest that these techniques often negatively affect the quality of the generated texts, and adversarial attacks can strip the watermarking signals, causing the texts to possibly evade detection. These findings have created resistance in the wide adoption of watermarking by LLM creators. Finally, to encourage adoption, we evaluate the robustness of several watermarking techniques to adversarial attacks by comparing paraphrasing and back translation (i.e., English $\to$ another language $\to$ English) attacks; and their ability to preserve quality and writing style of the unwatermarked texts by using linguistic metrics to capture quality and writing style of texts. Our results suggest that these watermarking techniques preserve semantics, deviate from the writing style of the unwatermarked texts, and are susceptible to adversarial attacks, especially for the back translation attack.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。