arXiv:2504.12229cs.LGcs.CL2025-04

发现人类和模型会无意识模仿水印特征,导致检测误判。

Watermarking Needs Input Repetition Masking

  • 通过分析对话适应性,揭示文本生成者会无意中复制水印信号。
  • 在看似不可能的场景中,人类与模型仍表现出水印模仿现象。
  • 提醒水印系统需更长序列与更低误报率,以保障长期可靠性。

大型语言模型(LLMs)的快速发展引发了滥用风险,如传播虚假信息。为此出现了两类应对措施:基于机器学习的检测器用于判断文本是否为合成内容,以及基于水印的技术,通过细微标记实现生成文本的识别与溯源。然而,研究表明人类在对话中会从句法到词汇层面调整语言风格。由此推断,人类或未加水印的模型可能无意中模仿出类似大模型生成文本的特征,从而干扰检测的可靠性。本文研究了这种‘模仿’现象的程度,发现无论是人类还是模型,都会在多种情境下无意识地复制水印信号,包括在看似不合理的条件下。这一发现挑战了现有学术假设,表明为了确保长期水印系统的有效性,必须显著降低误报率,并采用更长的词元序列作为水印种子。

原文摘要 · Abstract (English)

Recent advancements in Large Language Models (LLMs) raised concerns over potential misuse, such as for spreading misinformation. In response two counter measures emerged: machine learning-based detectors that predict if text is synthetic, and LLM watermarking, which subtly marks generated text for identification and attribution. Meanwhile, humans are known to adjust language to their conversational partners both syntactically and lexically. By implication, it is possible that humans or unwatermarked LLMs could unintentionally mimic properties of LLM generated text, making counter measures unreliable. In this work we investigate the extent to which such conversational adaptation happens. We call the concept $\textit{mimicry}$ and demonstrate that both humans and LLMs end up mimicking, including the watermarking signal even in seemingly improbable settings. This challenges current academic assumptions and suggests that for long-term watermarking to be reliable, the likelihood of false positives needs to be significantly lower, while longer word sequences should be used for seeding watermarking mechanisms.

水印技术大模型安全对抗检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。