提出高效估计混合文本中水印比例的新方法,解决真实场景下的生成内容溯源难题。
Optimal Estimation of Watermark Proportions in Hybrid AI-Human Texts
- 基于关键统计量构建混合模型,实现水印比例的可识别估计
- 在连续关键统计量下,所提估计器达到最优下界,误差极低
- 适用于主流无偏水印方案,适合内容安全与真实性检测研究者
大型语言模型中的文本水印是区分合成文本与人工撰写内容的重要工具。现有研究多关注整篇文本是否含水印,但现实场景常涉及人类与水印文本混合的内容。本文针对混合文本中水印比例的最优估计问题展开研究,将该问题建模为基于关键统计量的混合模型参数估计。我们证明,在某些水印方案下,该比例参数根本不可识别;而在采用连续关键统计量的水印方法中,仅需温和条件即可实现可识别性。针对此类方法,我们提出了高效的估计器,涵盖多种常见无偏水印方案,并推导出基于关键统计量的任意可测估计器的极小极大下界,证明所提估计器达到此下界。在合成数据及开源模型生成的混合文本上的实验表明,所提估计器始终具有高精度。
原文摘要 · Abstract (English)
Text watermarks in large language models (LLMs) are an increasingly important tool for detecting synthetic text and distinguishing human-written content from LLM-generated text. While most existing studies focus on determining whether entire texts are watermarked, many real-world scenarios involve mixed-source texts, which blend human-written and watermarked content. In this paper, we address the problem of optimally estimating the watermark proportion in mixed-source texts. We cast this problem as estimating the proportion parameter in a mixture model based on \emph{pivotal statistics}. First, we show that this parameter is not even identifiable in certain watermarking schemes, let alone consistently estimable. In stark contrast, for watermarking methods that employ continuous pivotal statistics for detection, we demonstrate that the proportion parameter is identifiable under mild conditions. We propose efficient estimators for this class of methods, which include several popular unbiased watermarks as examples, and derive minimax lower bounds for any measurable estimator based on pivotal statistics, showing that our estimators achieve these lower bounds. Through evaluations on both synthetic data and mixed-source text generated by open-source models, we demonstrate that our proposed estimators consistently achieve high estimation accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。