水印强度与生成效率不可兼得,论文揭示了二者根本矛盾。
Inevitable Trade-off between Watermark Strength and Speculative Sampling Efficiency for Language Models
- 证明水印强弱与采样效率无法同时达到最优
- 提出两种方法,各保其一,验证理论可行性
- 适合关注生成安全与推理加速的开发者参考
大语言模型是概率模型,内容生成本质是对其输出分布的采样。现有水印技术可在不改变输出质量的前提下注入水印,而加速技术如推测采样则利用草稿模型加快采样过程并保持输出分布不变。然而,目前尚无方法能同时实现加速与水印注入。本文研究此方向,发现二者集成非易事。我们证明了一个不可能定理:无法同时保持最强水印强度与最高采样效率。此外,我们提出了两种方法,分别维持采样效率或水印强度,但无法兼顾两者。本工作为理解大模型生成水印令牌时水印强度与采样效率之间的内在权衡提供了严格的理论基础,并通过数值实验验证了理论结果及所提方法的有效性。
原文摘要 · Abstract (English)
Large language models are probabilistic models, and the process of generating content is essentially sampling from the output distribution of the language model. Existing watermarking techniques inject watermarks into the generated content without altering the output quality. On the other hand, existing acceleration techniques, specifically speculative sampling, leverage a draft model to speed up the sampling process while preserving the output distribution. However, there is no known method to simultaneously accelerate the sampling process and inject watermarks into the generated content. In this paper, we investigate this direction and find that the integration of watermarking and acceleration is non-trivial. We prove a no-go theorem, which states that it is impossible to simultaneously maintain the highest watermark strength and the highest sampling efficiency. Furthermore, we propose two methods that maintain either the sampling efficiency or the watermark strength, but not both. Our work provides a rigorous theoretical foundation for understanding the inherent trade-off between watermark strength and sampling efficiency in accelerating the generation of watermarked tokens for large language models. We also conduct numerical experiments to validate our theoretical findings and demonstrate the effectiveness of the proposed methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。