arXiv:2510.24789cs.CLcs.CR2025-10被引 1

跨语言摘要能有效去除AI文本水印,且不损伤语义质量。

Cross-Lingual Summarization as a Black-Box Watermark Removal Attack

  • 将文本翻译到中间语言后摘要再回译,形成语义瓶颈破坏水印特征。
  • 在5种语言上测试,水印检测准确率降至接近随机水平(0.53 AUROC)。
  • 适合研究水印安全性的学者,也警示生成内容溯源的脆弱性。

水印被提出作为识别AI生成文本的轻量级机制,通常依赖对词元分布的扰动。尽管先前研究显示改写可削弱此类信号,但这些攻击仍部分可检测或损害文本质量。本文证明跨语言摘要攻击(CLSA)——即翻译至中间语言、摘要后再可选回译——构成更有效的攻击路径。通过跨语言施加语义瓶颈,CLSA系统性地消除词元级别的统计偏差,同时保持语义一致性。在多个水印方案(KGW、SIR、XSIR、Unigram)及五种语言(阿姆哈拉语、中文、印地语、西班牙语、斯瓦希里语)上的实验表明,与单语言改写相比,CLSA在相似文本质量下更显著降低水印检测准确率。在每语言300个保留样本上,CLSA持续使检测趋近随机水平。以专为跨语言鲁棒性设计的XSIR为例,改写时的AUROC为0.827,使用中文为中间语言的CWRA为0.823,而CLSA将其降至0.53(接近随机)。结果揭示了一条实用、低成本的跨语言水印移除路径,能压缩内容且无明显伪影。

原文摘要 · Abstract (English)

Watermarking has been proposed as a lightweight mechanism to identify AI-generated text, with schemes typically relying on perturbations to token distributions. While prior work shows that paraphrasing can weaken such signals, these attacks remain partially detectable or degrade text quality. We demonstrate that cross-lingual summarization attacks (CLSA) -- translation to a pivot language followed by summarization and optional back-translation -- constitute a qualitatively stronger attack vector. By forcing a semantic bottleneck across languages, CLSA systematically destroys token-level statistical biases while preserving semantic fidelity. In experiments across multiple watermarking schemes (KGW, SIR, XSIR, Unigram) and five languages (Amharic, Chinese, Hindi, Spanish, Swahili), we show that CLSA reduces watermark detection accuracy more effectively than monolingual paraphrase at similar quality levels. Our results highlight an underexplored vulnerability that challenges the practicality of watermarking for provenance or regulation. We argue that robust provenance solutions must move beyond distributional watermarking and incorporate cryptographic or model-attestation approaches. On 300 held-out samples per language, CLSA consistently drives detection toward chance while preserving task utility. Concretely, for XSIR (explicitly designed for cross-lingual robustness), AUROC with paraphrasing is $0.827$, with Cross-Lingual Watermark Removal Attacks (CWRA) [He et al., 2024] using Chinese as the pivot, it is $0.823$, whereas CLSA drives it down to $0.53$ (near chance). Results highlight a practical, low-cost removal pathway that crosses languages and compresses content without visible artifacts.

水印攻击跨语言文本摘要

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。