arXiv:2508.13131cs.CLcs.LG2025-08

通过融合水印与非水印检测,提升大模型生成内容的识别能力。

Improving Detection of Watermarked Language Models

  • 将水印检测与非水印检测方法结合,构建混合检测方案。
  • 在多种实验条件下,混合方案性能优于单一检测方式。
  • 适用于后训练模型(如指令微调、RLHF)中水印熵低的场景。

水印技术已成为识别大语言模型生成内容的有效手段。水印强度通常依赖于语言模型的熵以及输入提示集。然而,在实际应用中,熵值可能十分有限,特别是对于经过后训练(如指令微调或基于人类反馈的强化学习,RLHF)的大模型,仅依靠水印检测难以实现可靠识别。本文研究了融合水印检测器与非水印检测器是否能提升检测效果。我们探索了多种混合检测方案,在广泛实验条件下均观察到性能提升,显著优于单一检测类型。

原文摘要 · Abstract (English)

Watermarking has recently emerged as an effective strategy for detecting the generations of large language models (LLMs). The strength of a watermark typically depends strongly on the entropy afforded by the language model and the set of input prompts. However, entropy can be quite limited in practice, especially for models that are post-trained, for example via instruction tuning or reinforcement learning from human feedback (RLHF), which makes detection based on watermarking alone challenging. In this work, we investigate whether detection can be improved by combining watermark detectors with non-watermark ones. We explore a number of hybrid schemes that combine the two, observing performance gains over either class of detector under a wide range of experimental conditions.

水印检测大模型安全混合检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。