arXiv:2506.21602cs.CLcs.AI2025-06ICML被引 20

提出BiMark框架,实现高质量文本生成下的多比特水印嵌入与无模型检测。

BiMark: Unbiased Multilayer Watermarking for Large Language Models

  • 通过比特翻转重加权机制实现跨模型水印检测
  • 多层架构提升水印可检测性,短文本提取率高30%
  • 支持多比特水印编码,适合实际部署场景

大语言模型(LLM)的快速发展引发了对其生成内容真实性的担忧,监管机构亟需可靠的识别机制。尽管水印技术提供了可行方案,但现有方法难以同时满足文本质量保持、无模型检测和信息嵌入容量三大关键需求,核心挑战在于如何平衡文本质量与信息嵌入能力。为此,我们提出BiMark,一种新型水印框架,包含三项创新:(1) 比特翻转无偏重加权机制,实现无模型检测;(2) 多层架构,在不损害生成质量的前提下增强可检测性;(3) 信息编码方法,支持多比特水印。理论分析与大量实验表明,相较于当前最优多比特水印方法,BiMark在短文本上提取率最高提升30%,且困惑度更低,下游任务如摘要和翻译表现接近未加水印文本。

原文摘要 · Abstract (English)

Recent advances in Large Language Models (LLMs) have raised urgent concerns about LLM-generated text authenticity, prompting regulatory demands for reliable identification mechanisms. Although watermarking offers a promising solution, existing approaches struggle to simultaneously achieve three critical requirements: text quality preservation, model-agnostic detection, and message embedding capacity, which are crucial for practical implementation. To achieve these goals, the key challenge lies in balancing the trade-off between text quality preservation and message embedding capacity. To address this challenge, we propose BiMark, a novel watermarking framework that achieves these requirements through three key innovations: (1) a bit-flip unbiased reweighting mechanism enabling model-agnostic detection, (2) a multilayer architecture enhancing detectability without compromising generation quality, and (3) an information encoding approach supporting multi-bit watermarking. Through theoretical analysis and extensive experiments, we validate that, compared to state-of-the-art multi-bit watermarking methods, BiMark achieves up to 30% higher extraction rates for short texts while maintaining text quality indicated by lower perplexity, and performs comparably to non-watermarked text on downstream tasks such as summarization and translation.

水印LLM安全多比特

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。