比较两种方法估算文本中水印比例的样本效率,发现简化观测会损失效率。
Sample Complexities of Estimating Gumbel--Max Watermark Proportions with and without Reduction to Pivotal Statistics
- 用拉盖尔多项式和事件计数法分别处理简化与完整观测数据
- 简化观测下样本复杂度更高,完整观测可显著减少所需样本量
- 适合关注水印检测效率的研究者,尤其在资源受限场景
水印技术可实现大语言模型使用行为的统计可追溯,但真实文档通常混合了人工与模型生成内容。这引出了超越检测的量化问题:如何估计文档中来自特定水印模型的比例?本文研究在Gumbel--max水印机制下该比例的估计问题,将下一词预测分布视为未知且任意的干扰参数,仅满足非退化条件。比较两种观测模式:全观测模式下,估计器可获取每个位置的伪随机向量与选中的词;而在更常见的枢轴简化模式中,仅能观测一个服从一维均匀-贝塔混合分布的标量枢轴。在枢轴简化模式下,提出拉盖尔多项式估计器并建立匹配的信息论下界;在全观测模式下引入事件计数估计器,并证明匹配下界,表明其样本复杂度显著更低。结果表明,尽管枢轴简化是一种优雅且普遍的做法,但在估计水印比例时并不总是样本高效的。
原文摘要 · Abstract (English)
Watermarking promises statistical traceability of large language model (LLM) uses, but real documents rarely arrive as purely human-written or purely LLM-generated. This motivates a quantitative question beyond detection: what proportion of a document is generated from a pre-specified watermarked LLM? We study this watermark proportion estimation problem under the Gumbel--max watermarking mechanism, treating the next-token prediction distributions as unknown and arbitrary nuisance parameters subject to a non-degeneracy condition. We compare two observation regimes: in the full observation regime, the estimator observes the pseudorandom vector and the selected token at each position; in the more prevalent setting of pivotal reduction, it observes only a scalar pivot, which follows a one-dimensional Uniform--Beta mixture distribution. Under pivotal reduction, we develop a Laguerre-polynomial estimator and establish a matching information-theoretic lower bound for the sample complexity. For full observation, we introduce an event-counting estimator and show a matching lower bound, yielding a substantially smaller sample complexity. As our results imply, although reducing to pivotal statistics is an elegant and prevalent choice, it is not always sample-efficient for estimating the proportion of watermarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。