arXiv:2608.12713cs.CRcs.AI2026-08

一种可同时追踪来源并检测篡改的新型语言模型水印技术

Tracing Provenance and Detecting Tampering with Complementary LLM Watermarks

论文配图:Tracing Provenance and Detecting Tampering with Complementary LLM Watermarks
图 1 · 摘自论文原文
  • 通过共嵌入鲁棒与脆弱信号实现双重验证
  • 在两个模型和数据集上检测篡改率最高
  • 适合需要内容安全性的生成系统使用

为追踪大模型生成文本的来源,现有水印技术虽能抵抗编辑,但易被恶意篡改后仍保留归属,即“搭便车伪造”漏洞。本文提出一种新型水印机制,将鲁棒信号与脆弱信号共嵌入每个生成词元中:两者共享生成机制但使用独立密钥和不同的归一化文本种子窗口,前者抗编辑,后者对可见修改敏感。通过多轮无偏淘汰重加权保持生成分布,周期性分配策略调控两信号平衡。检测时依据双维评分区分‘完整’、‘篡改’、‘无水印’三种状态。在两个大模型及两个提示数据集上,该方法在篡改检测率上表现最优,同时保持良好的归属鲁棒性与困惑度。消融实验表明,可靠三态判断需明确定义完整状态、信号共嵌入及互补敏感性。

原文摘要 · Abstract (English)

Watermarking LLM-generated text is an important task for tracing its provenance. Existing LLM watermarks preserve provenance under editing, but this same robustness allows an adversary to alter critical content while retaining attribution, a vulnerability known as piggyback spoofing. We introduce an innovative watermark that jointly provides provenance and tamper evidence. It co-embeds a robust signal and a fragile signal into each generated token. The signals share the same mechanism but use independent keys and different seeding windows over normalized text, making one resilient to edits and the other sensitive to reader-visible changes. Multiple rounds of unbiased tournament reweighting preserve the expected generation distribution, while a periodic round-allocation pattern controls the trade-off between the two signals. At detection, their scores form a two-dimensional space supporting three decisions: Intact, Tampered, and No-Watermark. Across two large language models and two prompt datasets, our method demonstrates the highest tamper-detection rate among the evaluated methods while maintaining competitive attribution robustness and perplexity. Ablation studies show that reliable three-state detection requires a well-defined notion of intactness, co-embedding of the two signals, and complementary sensitivity to edits.

水印技术内容溯源对抗篡改LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。