破解大模型输出的水印技术,揭示现有方案存在可被轻易移除的漏洞。
Vaporizer: Breaking Watermarking Schemes for Large Language Model Outputs

- 通过改写、翻译和神经重写等攻击方式,系统性测试水印鲁棒性。
- 多种水印方案均能被有效去除,且语义保持基本不变。
- 适合关注生成内容安全与可信验证的研究者与开发者参考。
本文研究了当前最先进的大语言模型输出水印技术。这些方法声称具备鲁棒性、可扩展性和生产可用性,旨在促进大模型的负责任使用。我们针对一系列修改文本的攻击进行了分析,这些攻击在不改变文本整体语义的前提下进行有目标的语义调整。攻击策略包括词汇替换、机器翻译以及神经网络重写。攻击效果通过两个标准衡量:水印成功移除和语义内容保留。语义保留程度通过BERT得分、文本复杂度指标、语法错误率及Flesch阅读易度指数评估。实验结果表明,不同水印模型的有效性各异,但共同点是均可在合理努力下被移除。本研究揭示了现有水印系统的优缺点,为提升其安全性提供了改进方向。
原文摘要 · Abstract (English)
In this paper, we investigate the recent state-of-the-art schemes for watermarking large language models (LLMs) outputs. These techniques are claimed to be robust, scalable and production-grade, aimed at promoting responsible usage of LLMs. We analyse the effectiveness of these watermarking techniques against an extensive collection of modified text attacks, which perform targeted semantic changes without altering the general meaning of the text content. Our approach encompasses multiple attack strategies, which include lexical alterations, machine translation, and even neural paraphrasing. The attack efficacy is measured with two target criteria - successful removal of the watermark and preservation of semantic content. We evaluate semantic preservation through BERT scores, text complexity measures, grammatical errors, and Flesch Reading Ease indices. The experimental results reveal varying levels of effectiveness among different watermarking models, with the same underlying result that it is possible to remove the watermark with reasonable effort. This study sheds light on the strengths and weaknesses of existing LLM watermarking systems, suggesting how they should be constructed to improve security of available schemes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。