arXiv:2601.04534cs.CLcs.AI2026-01被引 1

针对孟加拉语大模型文本生成,提出分层水印方案提升抗翻译攻击能力。

BanglaLorica: Design and Evaluation of a Robust Watermarking Algorithm for Large Language Models in Bangla Text Generation

  • 采用嵌入层与生成后双重水印结合策略,增强鲁棒性。
  • 跨语言翻译攻击下检测准确率提升至40%-50%,比单层方法高3-4倍。
  • 适用于低资源语言,无需训练,适合版权保护与滥用检测场景。

随着大语言模型在文本生成中的广泛应用,水印技术成为作者溯源、知识产权保护和滥用检测的关键手段。尽管现有水印方法在高资源语言中表现良好,但其在低资源语言中的鲁棒性仍缺乏系统研究。本文首次对主流文本水印方法——KGW、指数采样(EXP)和Waterfall——在孟加拉语大模型文本生成中,面对跨语言往返翻译(RTT)攻击的性能进行了系统评估。在无攻击情况下,KGW与EXP检测准确率超过88%,且困惑度与ROUGE指标下降可忽略。然而,经RTT攻击后,检测准确率骤降至9%-13%,暴露出令牌级水印的根本性脆弱。为此,我们提出一种分层水印策略,融合嵌入时与生成后水印。实验表明,该策略使经RTT后的检测准确率提升25%-35%,达到40%-50%,相较单层方法相对提升3至4倍,仅带来可控的语义退化。研究量化了多语言水印中的鲁棒性-质量权衡,确立分层水印为低资源语言如孟加拉语的实用、免训练解决方案。代码与数据将公开。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly deployed for text generation, watermarking has become essential for authorship attribution, intellectual property protection, and misuse detection. While existing watermarking methods perform well in high-resource languages, their robustness in low-resource languages remains underexplored. This work presents the first systematic evaluation of state-of-the-art text watermarking methods: KGW, Exponential Sampling (EXP), and Waterfall, for Bangla LLM text generation under cross-lingual round-trip translation (RTT) attacks. Under benign conditions, KGW and EXP achieve high detection accuracy (>88%) with negligible perplexity and ROUGE degradation. However, RTT causes detection accuracy to collapse below RTT causes detection accuracy to collapse to 9-13%, indicating a fundamental failure of token-level watermarking. To address this, we propose a layered watermarking strategy that combines embedding-time and post-generation watermarks. Experimental results show that layered watermarking improves post-RTT detection accuracy by 25-35%, achieving 40-50% accuracy, representing a 3$\times$ to 4$\times$ relative improvement over single-layer methods, at the cost of controlled semantic degradation. Our findings quantify the robustness-quality trade-off in multilingual watermarking and establish layered watermarking as a practical, training-free solution for low-resource languages such as Bangla. Our code and data will be made public.

水印技术大模型低资源语言文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。