提出可随时检测的统计水印框架,提升生成内容识别效率
Towards Anytime-Valid Statistical Watermarking
- 基于e-value构建可随时检验的水印检测机制
- 相比现有方法,平均减少13-15%的检测令牌开销
- 适合需要高效、实时内容溯源的应用场景
大语言模型的普及催生了对机器生成文本识别的需求。尽管统计水印成为有前景的解决方案,但现有方法存在两大缺陷:采样分布选择缺乏理论依据,且依赖固定时长的假设检验,无法实现有效提前停止。本文提出首个基于e-value的水印框架——锚定e水印(Anchored E-Watermarking),将最优采样与任意时间有效的推断统一起来。不同于传统方法在可选停止下会破坏第一类错误控制,本框架通过构建检测过程的检验超鞅,实现有效、任意时间推断。通过引入锚定分布近似目标模型,我们刻画了最坏情况对数增长率下的最优e-value,并推导出最优期望停止时间。理论分析经仿真和主流基准测试验证,表明该框架可显著提升样本效率,使检测平均令牌预算降低13-15%相对于当前最优基线。
原文摘要 · Abstract (English)
The proliferation of Large Language Models (LLMs) necessitates efficient mechanisms to distinguish machine-generated content from human text. While statistical watermarking has emerged as a promising solution, existing methods suffer from two critical limitations: the lack of a principled approach for selecting sampling distributions and the reliance on fixed-horizon hypothesis testing, which precludes valid early stopping. In this paper, we bridge this gap by developing the first e-value-based watermarking framework, Anchored E-Watermarking, that unifies optimal sampling with anytime-valid inference. Unlike traditional approaches where optional stopping invalidates Type-I error guarantees, our framework enables valid, anytime-inference by constructing a test supermartingale for the detection process. By leveraging an anchor distribution to approximate the target model, we characterize the optimal e-value with respect to the worst-case log-growth rate and derive the optimal expected stopping time. Our theoretical claims are substantiated by simulations and evaluations on established benchmarks, showing that our framework can significantly enhance sample efficiency, reducing the average token budget required for detection by 13-15% relative to state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。