提出新型多比特水印框架,提升文本质量与鲁棒性
WorldCup Sampling for Multi-bit LLM Watermarking
- 将采样过程建模为通信信道,通过分层竞争机制嵌入信息
- 支持高容量水印,保持生成文本质量且能可靠恢复
- 适合需要可信溯源的生成式AI应用,如内容认证
随着大语言模型生成文本越来越接近人类水平,水印技术成为实现可靠归属的重要手段。多比特水印可编码更丰富的来源信息,但现有方法多基于静态logit扰动和计数解码,随载荷增加易降低文本质量并影响解码鲁棒性。本文提出WorldCup框架,将采样过程视为结构化通信信道,通过互补信号引导的分层竞争机制嵌入消息比特,并引入熵感知调制以保持生成质量,采用置信度感知解码实现鲁棒消息恢复。全面实验表明,WorldCup在消息容量、可检测性、鲁棒性、文本质量和解码效率间达到优异平衡,持续优于已有基线。本工作为多比特水印研究提供了可扩展且原理清晰的基础。
原文摘要 · Abstract (English)
As large language models (LLMs) generate increasingly human-like text, watermarking has emerged as a promising solution for reliable attribution beyond mere detection. While multi-bit watermarking enables richer provenance encoding, existing approaches typically extend zero-bit watermarking schemes by introducing static logit perturbations and counting-based decoding strategies, which can degrade text quality and compromise decoding robustness as the payload increases. In this paper, we propose WorldCup, a multi-bit watermarking framework for LLMs that models the sampling process as a structured communication channel and embeds message bits through a hierarchical competition mechanism guided by complementary signals. Moreover, WorldCup incorporates entropy-aware modulation to preserve generation quality and enables robust message recovery via confidence-aware decoding that accounts for token-level reliability. Comprehensive experiments demonstrate that WorldCup achieves a strong balance across message capacity, detectability, robustness, text quality, and decoding efficiency, consistently outperforming prior baselines. We believe that this work establishes a scalable and principled foundation for future research on multi-bit watermarking in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。