arXiv:2410.03168cs.CRcs.CL2024-10ICLR被引 17

提出通过精心设计提示识别水印模型,揭示当前水印易被用户发现。

Can Watermarked LLMs be Identified by Users via Crafted Prompts?

  • 设计Water-Probe算法,用特定提示探测水印存在
  • 多数主流水印算法在实验中可被轻易识别
  • 建议随机化密钥选择提升隐蔽性,提出Water-Bag策略

大型语言模型(LLM)文本水印技术在检测生成内容和防止滥用方面取得显著进展,现有方法具备高可检测性、低质量影响和对编辑的鲁棒性。然而,当前研究缺乏对水印在实际服务中不可察觉性的考察。这对模型提供者尤为重要,因披露水印可能降低用户使用意愿并增加被攻击风险。本文首次系统研究水印的不可察觉性,提出Water-Probe识别算法,通过精心设计的提示探测水印。核心思想是:相同水印密钥下,模型输出呈现一致偏差,不同密钥间差异相似。实验表明,几乎所有主流水印算法均能被该方法轻易识别,而Water-Probe对非水印模型误报率极低。最后,我们提出提升不可察觉性的关键在于增强水印密钥选择的随机性,并引入Water-Bag策略,通过融合多个密钥显著提升隐蔽性。

原文摘要 · Abstract (English)

Text watermarking for Large Language Models (LLMs) has made significant progress in detecting LLM outputs and preventing misuse. Current watermarking techniques offer high detectability, minimal impact on text quality, and robustness to text editing. However, current researches lack investigation into the imperceptibility of watermarking techniques in LLM services. This is crucial as LLM providers may not want to disclose the presence of watermarks in real-world scenarios, as it could reduce user willingness to use the service and make watermarks more vulnerable to attacks. This work is the first to investigate the imperceptibility of watermarked LLMs. We design an identification algorithm called Water-Probe that detects watermarks through well-designed prompts to the LLM. Our key motivation is that current watermarked LLMs expose consistent biases under the same watermark key, resulting in similar differences across prompts under different watermark keys. Experiments show that almost all mainstream watermarking algorithms are easily identified with our well-designed prompts, while Water-Probe demonstrates a minimal false positive rate for non-watermarked LLMs. Finally, we propose that the key to enhancing the imperceptibility of watermarked LLMs is to increase the randomness of watermark key selection. Based on this, we introduce the Water-Bag strategy, which significantly improves watermark imperceptibility by merging multiple watermark keys.

水印检测LLM安全提示工程隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。