arXiv:2501.12174cs.LG2025-01被引 5

用双极水印提升大模型文本水印检测精度,无需额外计算资源。

BiMarker: Enhancing Text Watermark Detection for Large Language Models with Bipolar Watermarks

  • 将生成文本分为正负两极,增强水印可检测性
  • 在不增加计算成本下显著提升水印识别率
  • 适合需要高保真水印检测的AI内容监管场景

大型语言模型(LLMs)的快速发展引发了对区分人工智能生成文本与人类内容的担忧。现有水印技术如\kgw面临水印强度低和误报率严格的问题。我们的分析表明,当前方法依赖对非水印文本的粗略估计,限制了水印可检测性。为此,我们提出双极水印(\tool),将生成文本划分为正负两极,可在不增加额外计算资源或提示信息的前提下提升检测效果。理论分析与实验结果表明,\tool具有有效性且兼容现有优化技术,为大模型生成内容的水印提供新的优化维度。

原文摘要 · Abstract (English)

The rapid growth of Large Language Models (LLMs) raises concerns about distinguishing AI-generated text from human content. Existing watermarking techniques, like \kgw, struggle with low watermark strength and stringent false-positive requirements. Our analysis reveals that current methods rely on coarse estimates of non-watermarked text, limiting watermark detectability. To address this, we propose Bipolar Watermark (\tool), which splits generated text into positive and negative poles, enhancing detection without requiring additional computational resources or knowledge of the prompt. Theoretical analysis and experimental results demonstrate \tool's effectiveness and compatibility with existing optimization techniques, providing a new optimization dimension for watermarking in LLM-generated content.

水印检测大模型文本安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。