arXiv:2510.01637cs.LG2025-10被引 2

通过组合水印技术检测大模型生成文本的后期修改

Detecting Post-generation Edits to Watermarked LLM Outputs via Combinatorial Watermarking

  • 用词汇子集的确定性组合模式嵌入水印
  • 可定位修改区域,错误率低于5%且准确率达90%以上
  • 适合需要内容溯源与防篡改的AI生成场景

水印已成为专有语言模型的关键技术,用于区分人工智能生成与人工撰写的文本。然而在许多实际场景中,大型语言模型生成的内容可能经历后期编辑,如人工修订或伪造攻击,因此检测并定位这些修改至关重要。本文提出一项新任务:检测对水印化大模型输出的局部后期编辑。为此,我们设计了一种基于组合模式的水印框架,将词汇表划分为不相交子集,并在生成过程中强制执行确定性的组合模式以嵌入水印。同时,我们引入全局统计量用于水印检测,并设计轻量级局部统计量来标记和定位潜在修改。我们提出了两种任务特定评估指标——第一类错误率和检测准确率,在多种编辑场景下对开源大模型进行评估,结果表明该方法在编辑定位上具有优异的实证性能。

原文摘要 · Abstract (English)

Watermarking has become a key technique for proprietary language models, enabling the distinction between AI-generated and human-written text. However, in many real-world scenarios, LLM-generated content may undergo post-generation edits, such as human revisions or even spoofing attacks, making it critical to detect and localize such modifications. In this work, we introduce a new task: detecting post-generation edits locally made to watermarked LLM outputs. To this end, we propose a combinatorial pattern-based watermarking framework, which partitions the vocabulary into disjoint subsets and embeds the watermark by enforcing a deterministic combinatorial pattern over these subsets during generation. We accompany the combinatorial watermark with a global statistic that can be used to detect the watermark. Furthermore, we design lightweight local statistics to flag and localize potential edits. We introduce two task-specific evaluation metrics, Type-I error rate and detection accuracy, and evaluate our method on open-source LLMs across a variety of editing scenarios, demonstrating strong empirical performance in edit localization.

水印技术文本溯源大模型安全生成检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。