arXiv:2511.06645stat.MLcs.CL2025-11中稿 · publication in STA…被引 1

提出自适应方法,精准分割文本中的水印与非水印段落。

Adaptive Testing for Segmenting Watermarked Texts From Language Models

  • 基于似然加权与逆变换采样,改进水印检测框架。
  • 无需精确提示词估计,可准确识别混合文本中的水印片段。
  • 适用于教育防伪、信息可信度验证等场景。

大型语言模型(如 GPT-4、Claude 3.5)的广泛应用凸显了区分模型生成文本与人工撰写内容的重要性,以遏制虚假信息传播及教育滥用。一种有前景的方法是水印技术,通过在生成文本中嵌入细微的统计信号实现可靠识别。本文首先推广了先前研究的基于似然的检测方法,引入灵活的加权形式,并进一步将其适配至逆变换采样方法。超越单一检测,我们还将该自适应策略扩展至更复杂的任务——将给定文本分割为水印与非水印子串。与依赖高敏感性提示词估计的现有方法不同,本框架无需精确提示词估计。大量数值实验表明,该方法在识别含混合水印与非水印内容的文本时兼具高效性与鲁棒性。

原文摘要 · Abstract (English)

The rapid adoption of large language models (LLMs), such as GPT-4 and Claude 3.5, underscores the need to distinguish LLM-generated text from human-written content to mitigate the spread of misinformation and misuse in education. One promising approach to address this issue is the watermark technique, which embeds subtle statistical signals into LLM-generated text to enable reliable identification. In this paper, we first generalize the likelihood-based LLM detection method of a previous study by introducing a flexible weighted formulation, and further adapt this approach to the inverse transform sampling method. Moving beyond watermark detection, we extend this adaptive detection strategy to tackle the more challenging problem of segmenting a given text into watermarked and non-watermarked substrings. In contrast to the approach in a previous study, which relies on accurate estimation of next-token probabilities that are highly sensitive to prompt estimation, our proposed framework removes the need for precise prompt estimation. Extensive numerical experiments demonstrate that the proposed methodology is both effective and robust in accurately segmenting texts containing a mixture of watermarked and non-watermarked content.

水印检测文本分割LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。