提出可定位混合来源文本中水印位置的自适应方法,提升生成内容溯源能力。
Optimal Watermark Localization in Mixed-Source Large Language Model Texts

- 基于关键统计量构建分词级多重检验模型,识别水印残留位置。
- 理论证明发现水印比检测更难,且在特定参数下无法一致分类。
- 无需先验知识,通过数据驱动估计水印存活比例,性能接近最优。
水印为大语言模型生成文本的认证提供了有效途径。然而实际中,文本经重写、插入、删除或改写后,水印信号仅在部分词元位置残留。尽管已有研究关注水印全局检测,但其定位仍不明确。本文将水印定位建模为基于关键统计量的词元级多重检验问题,引入隐变量记录每个位置水印依赖是否留存。在由信号稀疏度、下一词集中度和有效词表增长指数定义的渐近框架下,推导出全局检测的精确边界,以及坐标系基定位规则下的发现与分类相变现象。结果表明,发现严格难于检测,且在此类方法中无法实现一致分类。随后提出一种自适应阈值方法,无需知晓指数或时变的下一词分布,仅需数据驱动估计残留水印比例,即可达到最优发现边界,并具有接近同质基定位规则的发现效能。模拟验证了理论相变行为,模型生成文本实验也展示了在常见编辑机制下的实用定位性能。
原文摘要 · Abstract (English)
Watermarking provides a principled way to authenticate text generated by large language models (LLMs). In practice, however, the final text may be mixed-source, with watermark evidence surviving at only a subset of token positions after rewriting, insertion, deletion, or paraphrasing. Although prior work has studied global detection of watermark signals, when such signals can be localized remains unclear. We formulate watermark localization as a token-level multiple-testing problem based on pivotal statistics, with a latent indicator recording whether watermark dependence survives at each position. Under an asymptotic regime indexed by exponents for signal sparsity, next-token concentration, and effective-vocabulary growth, we derive a sharp boundary for global detection and phase transitions for discovery and classification within the class of coordinatewise pivot-based localization rules. We show that discovery is strictly harder than detection and that consistent classification is impossible across the parameter regime within this class. We then develop an adaptive thresholding method that does not require knowledge of the exponents or time-varying next-token distributions, but uses a data-driven estimate of the surviving watermark fraction. The method attains the optimal discovery boundary and near-optimal discovery power relative to homogeneous pivot-based rules. Simulations support the theoretical phase transitions, while experiments on model-generated texts demonstrate practical localization performance under common edit mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。