评测大模型文本水印在攻击下的鲁棒性,揭示设计选择的影响。
Watermark under Fire: A Robustness Evaluation of LLM Watermarking
- 构建统一平台WaterPark,集成10种水印方法与12种攻击手段。
- 实验发现水印鲁棒性差异显著,部分设计可提升抗攻击能力。
- 为对抗环境中的水印部署提供最优实践建议。
针对大模型生成文本的水印技术,现有方法缺乏统一评估平台,诸多关键问题仍待探索:水印方案的优劣与抗攻击能力如何?不同设计选择对鲁棒性有何影响?在对抗环境中如何最优使用?为此,本文系统梳理了现有水印方法与移除攻击,构建出涵盖10种前沿水印技术与12类典型攻击的WaterPark统一平台。通过该平台,我们全面评估了现有水印方案的抗攻击表现,揭示了设计因素对其鲁棒性的影响。同时,研究提出了对抗环境下水印部署的最佳实践。本工作不仅深化了对当前水印技术的理解,也为未来研究提供了可复现的测试基准。
原文摘要 · Abstract (English)
Various watermarking methods (``watermarkers'') have been proposed to identify LLM-generated texts; yet, due to the lack of unified evaluation platforms, many critical questions remain under-explored: i) What are the strengths/limitations of various watermarkers, especially their attack robustness? ii) How do various design choices impact their robustness? iii) How to optimally operate watermarkers in adversarial environments? To fill this gap, we systematize existing LLM watermarkers and watermark removal attacks, mapping out their design spaces. We then develop WaterPark, a unified platform that integrates 10 state-of-the-art watermarkers and 12 representative attacks. More importantly, by leveraging WaterPark, we conduct a comprehensive assessment of existing watermarkers, unveiling the impact of various design choices on their attack robustness. We further explore the best practices to operate watermarkers in adversarial environments. We believe our study sheds light on current LLM watermarking techniques while WaterPark serves as a valuable testbed to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。