用经典统计检验提升大模型文本水印的检测力与鲁棒性
On the Empirical Power of Goodness-of-Fit Tests in Watermark Detection
- 采用八种拟合优度检验,评估水印检测性能
- 低温度生成时的重复文本让检验更具优势
- 适合关注内容溯源与防伪的开发者
大语言模型(LLMs)可大规模生成类人文本,引发内容真实性的担忧。文本水印通过在生成文本中嵌入可检测的统计信号,提供内容来源验证的可信手段。许多检测方法依赖于人类写作文本下独立同分布(i.i.d.)的枢轴统计量,使拟合优度(GoF)检验成为水印检测的自然工具。然而,该方法在此场景仍鲜被研究。本文系统评估了八种GoF检验在三种主流水印方案中的表现,使用三个开源大模型、两个数据集、多种生成温度及多重后编辑方法。结果表明,通用的GoF检验能显著提升水印检测的检测力与鲁棒性。尤其发现,低温度生成中常见的文本重复现象,为GoF检验提供了现有方法未利用的独特优势。研究强调,经典GoF检验是简单但强大且被低估的水印检测工具。
原文摘要 · Abstract (English)
Large language models (LLMs) raise concerns about content authenticity and integrity because they can generate human-like text at scale. Text watermarks, which embed detectable statistical signals into generated text, offer a provable way to verify content origin. Many detection methods rely on pivotal statistics that are i.i.d. under human-written text, making goodness-of-fit (GoF) tests a natural tool for watermark detection. However, GoF tests remain largely underexplored in this setting. In this paper, we systematically evaluate eight GoF tests across three popular watermarking schemes, using three open-source LLMs, two datasets, various generation temperatures, and multiple post-editing methods. We find that general GoF tests can improve both the detection power and robustness of watermark detectors. Notably, we observe that text repetition, common in low-temperature settings, gives GoF tests a unique advantage not exploited by existing methods. Our results highlight that classic GoF tests are a simple yet powerful and underused tool for watermark detection in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。