arXiv:2512.16182cs.CRcs.CL2025-12ACL

DualGuard首次同时防御改写和伪造攻击,提升大模型水印可靠性。

DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack

  • 双流动态注入水印信号,根据语义内容自适应生成互补水印。
  • 在多个数据集上实现高检测率、强鲁棒性与可追溯性。
  • 适合需防伪造内容篡改的AI生成内容可信场景。

随着云服务快速发展,大语言模型通过各类网络平台日益普及,但也面临滥用风险。模型水印成为应对滥用、保护知识产权的有效手段。现有水印算法主要针对改写攻击,忽视了可注入有害内容、破坏水印可信度的搭便车伪造攻击。为此,本文提出DualGuard,首个能同时防御改写与伪造攻击的水印算法。其采用自适应双流水印机制,根据语义内容动态注入两种互补水印信号,实现对伪造攻击的检测与溯源,保障水印检测的可靠性和可信性。在多个数据集和语言模型上的实验表明,DualGuard在检测性、鲁棒性、可追溯性及文本质量方面均表现优异,显著推进了大模型水印在真实应用中的发展。

原文摘要 · Abstract (English)

With the rapid development of cloud-based services, large language models have become increasingly accessible through various web platforms. However, this accessibility has also led to growing risks of model abuse. LLM watermarking has emerged as an effective approach to mitigate such misuse and protect intellectual property. Existing watermarking algorithms, however, primarily focus on defending against paraphrase attacks while overlooking piggyback spoofing attacks, which can inject harmful content, compromise watermark reliability, and undermine trust in attribution. To address this limitation, we propose DualGuard, the first watermarking algorithm capable of defending against both paraphrase and spoofing attacks. DualGuard employs the adaptive dual-stream watermarking mechanism, in which two complementary watermark signals are dynamically injected based on the semantic content. This design enables DualGuard not only to detect but also to trace spoofing attacks, thereby ensuring reliable and trustworthy watermark detection. Extensive experiments conducted across multiple datasets and language models demonstrate that DualGuard achieves excellent detectability, robustness, traceability, and text quality, effectively advancing the state of LLM watermarking for real-world applications.

大模型水印伪造攻击双流机制可信生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。