提升100多种语言的文本水印鲁棒性,解决低资源语言水印失效问题
Is Multilingual LLM Watermarking Truly Multilingual? Scaling Robustness to 100+ Languages via Back-Translation
- 通过反向翻译搜索最佳语言恢复水印强度
- 在中低资源语言上平均提升0.23 AUC、TPR@1%提高37%
- 兼容多种水印方法,无需修改模型,适合多语言场景
多语言水印旨在使大语言模型输出在跨语言间可追溯,但现有方法仍存不足。尽管宣称具备跨语言鲁棒性,其评估仅限高资源语言。我们发现,现有方法在中低资源语言的翻译攻击下无法保持鲁棒性,根源在于语义聚类失效——当分词器词汇表中该语言完整词项过少时。为此,我们提出STEAM,一种基于贝叶斯优化的方法,在133种候选语言中搜索最优反向翻译路径以恢复水印强度。该方法兼容任意水印技术,对分词器和语言均具鲁棒性,非侵入式且易扩展至新语言。平均实现+0.23 AUC与+37% TPR@1%提升,为实现更公平的多语言水印提供可扩展方案。
原文摘要 · Abstract (English)
Multilingual watermarking aims to make large language model (LLM) outputs traceable across languages, yet current methods still fall short. Despite claims of cross-lingual robustness, they are evaluated only on high-resource languages. We show that existing multilingual watermarking methods are not truly multilingual: they fail to remain robust under translation attacks in medium- and low-resource languages. We trace this failure to semantic clustering, which fails when the tokenizer vocabulary contains too few full-word tokens for a given language. To address this, we introduce STEAM, a detection method that uses Bayesian optimisation to search among 133 candidate languages for the back-translation that best recovers the watermark strength. It is compatible with any watermarking method, robust across different tokenizers and languages, non-invasive, and easily extendable to new languages. With average gains of +0.23 AUC and +37% TPR@1%, STEAM provides a scalable approach toward fairer watermarking across the diversity of languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。