arXiv:2510.22007math.STcs.CL2025-10被引 3

破解语言水印检测难题,实现更精准的模型输出追踪。

Optimal Detection for Language Watermarks with Pseudorandom Collision

  • 构建分层分区统计框架,识别最小独立单元以应对重复文本依赖
  • 提出非渐近效率度量,实现最优检测规则与严格第一类错误控制
  • 适用于Gumbel-max和逆变换水印,理论与实验均验证性能提升

文本水印在确保大语言模型输出可追溯性和防止滥用方面至关重要。然而,现有方法多假设伪随机性完美,实际生成文本中的重复会引发碰撞,产生结构化依赖,破坏第一类错误控制并使标准分析失效。本文提出一种统计框架,通过双层分层分区捕捉该结构,核心是‘最小单元’——组内允许依赖但组间独立的最小单位。基于最小单元定义非渐近效率度量,并将水印检测建模为极小极大假设检验问题。应用于Gumbel-max与逆变换水印,导出闭式最优检测规则。解释了为何舍弃重复统计量常能提升性能,并表明除非退化,否则必须处理组内依赖。理论与实验均证实,在严格控制第一类错误下检测能力显著提升。本工作首次为不完美伪随机性下的水印检测提供了严谨基础,兼具理论洞见与实践指导意义。

原文摘要 · Abstract (English)

Text watermarking plays a crucial role in ensuring the traceability and accountability of large language model (LLM) outputs and mitigating misuse. While promising, most existing methods assume perfect pseudorandomness. In practice, repetition in generated text induces collisions that create structured dependence, compromising Type I error control and invalidating standard analyses. We introduce a statistical framework that captures this structure through a hierarchical two-layer partition. At its core is the concept of minimal units -- the smallest groups treatable as independent across units while permitting dependence within. Using minimal units, we define a non-asymptotic efficiency measure and cast watermark detection as a minimax hypothesis testing problem. Applied to Gumbel-max and inverse-transform watermarks, our framework produces closed-form optimal rules. It explains why discarding repeated statistics often improves performance and shows that within-unit dependence must be addressed unless degenerate. Both theory and experiments confirm improved detection power with rigorous Type I error control. These results provide the first principled foundation for watermark detection under imperfect pseudorandomness, offering both theoretical insight and practical guidance for reliable tracing of model outputs.

语言水印统计检测大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。