arXiv:2605.00348cs.CRcs.CL2026-05中稿 · ICML被引 1

提出新水印框架,解决多比特水印误检率高的问题。

Block-wise Codeword Embedding for Reliable Multi-bit Text Watermarking

论文配图:Block-wise Codeword Embedding for Reliable Multi-bit Text Watermarking
图 1 · 摘自论文原文
  • 分块投票+窗口滑移验证,分离嵌入与检测逻辑
  • 10%同义替换下仍保持96.5%识别率、2%误检率
  • 适用于模型溯源,对大模型水印部署有重要意义

现有大语言模型多比特水印方法过度追求容量而忽视可靠性,常将解码与检测混淆。分析表明,基于纠错码的提取器存在灾难性高误检率(FPR),设置拒识阈值仅导致检测灵敏度(TPR)退化为随机猜测。为此,我们提出BREW(分块可靠嵌入水印框架),采用两阶段机制:(i) 独立分块投票实现盲消息估计;(ii) 窗口滑移验证严格校验载荷对局部编辑的鲁棒性。实验显示,在10%同义替换条件下,BREW实现TPR=0.965、FPR=0.02,证明高误检并非多比特水印固有缺陷,而是传统以解码为中心设计的结构性问题。该框架模型无关且理论完备,为可靠取证部署提供可扩展解决方案。

原文摘要 · Abstract (English)

Recent multi-bit watermarking methods for large language models (LLMs) prioritize capacity over reliability, often conflating decoding with detection. Our analysis reveals that existing ECC-based extractors suffer from catastrophic false positive rates (FPR), and applying rejection thresholds merely collapses detection sensitivity (TPR) to random guessing. To resolve this structural limitation, we propose BREW (Block-wise Reliable Embedding for Watermarking), a framework shifting the paradigm to designated verification. BREW employs a two-stage mechanism: (i) blind message estimation via independent block voting, followed by (ii) window-shifting verification that rigorously validates the payload against local edits. Experiments demonstrate that BREW achieves a TPR of 0.965 with an FPR of 0.02 under 10% synonym substitution, demonstrating that the high-FPR issue is not an inherent trade-off of multi-bit watermarking, but a solvable structural flaw of prior decoding-centric designs. Our framework is model-agnostic and theoretically grounded, providing a scalable solution for reliable forensic deployment.

水印技术大模型安全信息隐藏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。