提出统一评估框架,全面测试文本水印的实用性。
CEFW: A Comprehensive Evaluation Framework for Watermark in Large Language Models
- 构建五维评估体系,覆盖检测难易、文本质量等关键指标。
- 新方法BW在所有维度表现优于现有技术,尤其兼顾鲁棒与隐蔽性。
- 适合关注生成内容溯源与安全的AI研究者使用。
文本水印为识别大模型生成的合成文本提供了有效方案。然而,现有技术常只关注特定标准,忽视其他关键方面,缺乏统一评估。为此,我们提出综合水印评估框架(CEFW),从五个核心维度全面评估水印方法:检测难易度、文本质量保真度、嵌入开销最小化、对抗攻击下的鲁棒性,以及防止仿冒或伪造的隐蔽性。通过综合考量这些指标,CEFW可系统评估水印的实际效果与适用性。此外,我们提出一种简单高效的水印方法——平衡水印(BW),通过均衡信息嵌入方式确保鲁棒性与隐蔽性。大量实验表明,BW在所有评估维度上均优于现有方法。代码已开源,供后续研究使用。https://github.com/DrankXs/BalancedWatermark
原文摘要 · Abstract (English)
Text watermarking provides an effective solution for identifying synthetic text generated by large language models. However, existing techniques often focus on satisfying specific criteria while ignoring other key aspects, lacking a unified evaluation. To fill this gap, we propose the Comprehensive Evaluation Framework for Watermark (CEFW), a unified framework that comprehensively evaluates watermarking methods across five key dimensions: ease of detection, fidelity of text quality, minimal embedding cost, robustness to adversarial attacks, and imperceptibility to prevent imitation or forgery. By assessing watermarks according to all these key criteria, CEFW offers a thorough evaluation of their practicality and effectiveness. Moreover, we introduce a simple and effective watermarking method called Balanced Watermark (BW), which guarantees robustness and imperceptibility through balancing the way watermark information is added. Extensive experiments show that BW outperforms existing methods in overall performance across all evaluation dimensions. We release our code to the community for future research. https://github.com/DrankXs/BalancedWatermark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。