提出新方法Tr-GoF,有效识别经人工修改的AI生成文本水印。
Robust Detection of Watermarks for Large Language Models Under Human Edits
- 用截断拟合优度测试建模人类编辑,提升水印鲁棒性。
- 在大量修改和弱信号下,检测效率优于现有方法。
- 无需了解编辑程度或模型概率,适合实际应用。
水印技术为区分大语言模型(LLMs)生成文本与人工撰写文本提供了有效手段。然而,人工对LLM生成文本的频繁修改会削弱水印信号,显著降低现有检测方法的性能。本文通过混合模型检测建模人类编辑,提出一种截断拟合优度检验(Tr-GoF)方法,用于在人工编辑条件下检测水印文本。我们证明,在大量文本修改和水印信号趋近于零的渐近场景中,Tr-GoF在检测Gumbel-max水印时达到最优性。尤为重要的是,该方法能自适应实现最优性,无需精确掌握人工编辑水平或语言模型的概率特性,而传统最优但不实用的Neyman-Pearson似然比检验则依赖这些先验信息。此外,我们在中等文本修改场景下也证明了Tr-GoF具有最高检测效率。相反,现有方法采用的求和型检测规则因统计量的可加性,在两类场景下均无法实现最优鲁棒性。最后,我们在合成数据及OPT、LLaMA系列开源大模型上验证了Tr-GoF的竞争力,其性能有时更优。
原文摘要 · Abstract (English)
Watermarking has offered an effective approach to distinguishing text generated by large language models (LLMs) from human-written text. However, the pervasive presence of human edits on LLM-generated text dilutes watermark signals, thereby significantly degrading detection performance of existing methods. In this paper, by modeling human edits through mixture model detection, we introduce a new method in the form of a truncated goodness-of-fit test for detecting watermarked text under human edits, which we refer to as Tr-GoF. We prove that the Tr-GoF test achieves optimality in robust detection of the Gumbel-max watermark in a certain asymptotic regime of substantial text modifications and vanishing watermark signals. Importantly, Tr-GoF achieves this optimality \textit{adaptively} as it does not require precise knowledge of human edit levels or probabilistic specifications of the LLMs, in contrast to the optimal but impractical (Neyman--Pearson) likelihood ratio test. Moreover, we establish that the Tr-GoF test attains the highest detection efficiency rate in a certain regime of moderate text modifications. In stark contrast, we show that sum-based detection rules, as employed by existing methods, fail to achieve optimal robustness in both regimes because the additive nature of their statistics is less resilient to edit-induced noise. Finally, we demonstrate the competitive and sometimes superior empirical performance of the Tr-GoF test on both synthetic data and open-source LLMs in the OPT and LLaMA families.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。