arXiv:2602.01428cs.LGcs.CR2026-02中稿 · ICLR

提升语言模型水印强度的同时保持采样效率,打破二者冲突的困局。

Improving the Trade-off Between Watermark Strength and Speculative Sampling Efficiency for Language Models

  • 基于伪随机数构造水印,量化水印强度并建模其与采样效率的关系
  • 通过优化水印策略,在不降低采样效率前提下显著提升可检测性
  • 适用于需要高效且可追溯生成内容的部署场景,如AI安全与合规

水印是追踪大语言模型输出来源的可靠方法,但实际应用受推理效率限制。推测采样可加速推理,效率随草稿模型与目标模型的接受率提升而提高。然而,已有研究表明水印强度与接受率存在根本性权衡:强度越高,接受率越低,难以兼顾。本文重新审视该权衡,发现其并非绝对。提出一种衡量水印强度的定量指标,该指标在标记符为伪随机数确定性函数时达到最大,从而控制统计可检测性。基于此,将权衡问题建模为约束优化,并推导出两种现有水印方案的显式帕累托曲线。最后,提出一种原则性机制,在草稿令牌接受过程中注入伪随机性,实现水印强度最大化的同时维持推测采样效率。实验表明,该方法在不牺牲效率的前提下提升了可检测性。研究揭示了推测采样与水印统一的原则,推动二者在实际中的高效部署。

原文摘要 · Abstract (English)

Watermarking is a principled approach for tracing the provenance of large language model (LLM) outputs, but its deployment in practice is hindered by inference inefficiency. Speculative sampling accelerates inference, with efficiency improving as the acceptance rate between draft and target models increases. Yet recent work reveals a fundamental trade-off: higher watermark strength reduces acceptance, preventing their simultaneous achievement. We revisit this trade-off and show it is not absolute. We introduce a quantitative measure of watermark strength that governs statistical detectability and is maximized when tokens are deterministic functions of pseudorandom numbers. Using this measure, we fully characterize the trade-off as a constrained optimization problem and derive explicit Pareto curves for two existing watermarking schemes. Finally, we introduce a principled mechanism that injects pseudorandomness into draft-token acceptance, ensuring maximal watermark strength while maintaining speculative sampling efficiency. Experiments further show that this approach improves detectability without sacrificing efficiency. Our findings uncover a principle that unites speculative sampling and watermarking, paving the way for their efficient and practical deployment.

水印技术推理加速语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。