为大模型水印设计新框架,平衡文本质量与信息容量
Distributional Information Embedding: A Framework for Multi-bit Watermarking
- 通过调整生成分布嵌入信息,实现可检测的多比特水印
- 理论证明水印速率上限等于模型输出熵,可随允许失真提升
- 适用于对水印鲁棒性、误报率有严格要求的场景
本文提出一种新型问题——分布信息嵌入,源于大语言模型多比特水印的实际需求。与传统信息嵌入不同,大模型水印主动控制文本生成过程,通过调整词元分布来嵌入可检测信号。我们构建了信息论框架,分析文本质量、可检测性与信息速率三者间的根本权衡。在渐近情形下,证明当错误率趋近于零时,最大可实现速率等于模型输出分布的熵,且随允许失真增大而上升。同时刻画了达到该速率的最优水印方案。在有限长度非独立同分布词元情况下,识别出在假阳性率和失真约束下最大化检测概率的方案。
原文摘要 · Abstract (English)
This paper introduces a novel problem, distributional information embedding, motivated by the practical demands of multi-bit watermarking for large language models (LLMs). Unlike traditional information embedding, which embeds information into a pre-existing host signal, LLM watermarking actively controls the text generation process--adjusting the token distribution--to embed a detectable signal. We develop an information-theoretic framework to analyze this distributional information embedding problem, characterizing the fundamental trade-offs among three critical performance metrics: text quality, detectability, and information rate. In the asymptotic regime, we demonstrate that the maximum achievable rate with vanishing error corresponds to the entropy of the LLM's output distribution and increases with higher allowable distortion. We also characterize the optimal watermarking scheme to achieve this rate. Extending the analysis to the finite-token case with non-i.i.d. tokens, we identify schemes that maximize detection probability while adhering to constraints on false alarm and distortion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。