arXiv:2604.16923cs.AI2026-04被引 2

通过可证明的对齐印记,实现零样本高效检测AI生成文本。

Alignment Imprint: Zero-Shot AI-Generated Text Detection via Provable Preference Discrepancy

论文配图:Alignment Imprint: Zero-Shot AI-Generated Text Detection via Provable Preference Discrepancy
图 1 · 摘自论文原文
  • 从对齐过程推导出可度量的分布印记,作为检测依据。
  • 提出LAPD统计量,在所有场景下性能比现有方法提升45.82%。
  • 适合需要零样本、高稳定性的文本检测应用者。

检测AI生成文本是一项重要但具挑战性的问题。现有基于似然的方法常受内容复杂性影响,表现不稳定。本文的核心洞察是:现代大语言模型在对齐(包括微调和偏好对齐)过程中会留下可度量的分布印记。我们通过将对齐过程抽象为一系列约束优化步骤,理论上推导出该印记,发现对数似然比可自然分解为隐式指令偏差与偏好奖励。我们称此量为对齐印记。为缓解高熵区域的不稳定性,引入基于对齐印记的信息加权标准化统计量——对数似然对齐偏好差异(LAPD)。我们提供理论保证,表明基于对齐的统计量性能优于Fast-DetectGPT。同时理论证明,当对齐模型与基础模型分布接近时,LAPD严格优于未加权的对齐得分。大量实验显示,LAPD相对最强基线提升45.82%,在所有设置中均取得显著且一致的增益。

原文摘要 · Abstract (English)

Detecting AI-generated text is an important but challenging problem. Existing likelihood-based detection methods are often sensitive to content complexity and may exhibit unstable performance. In this paper, our key insight is that modern Large Language Models (LLMs) undergo alignment (including fine-tuning and preference tuning), leaving a measurable distributional imprint. We theoretically derive this imprint by abstracting the alignment process as a sequence of constrained optimization steps, showing that the log-likelihood ratio can naturally decompose into implicit instructional biases and preference rewards. We refer to this quantity as the Alignment Imprint. Furthermore, to mitigate the instability in high-entropy regions, we introduce Log-likelihood Alignment Preference Discrepancy (LAPD), a standardized information-weighted statistic based on alignment imprint. We provide statistical guarantee that alignment-based statistics dominate Fast-DetectGPT in performance. We also theoretically show that LAPD strictly improves the unweighted alignment scores when the aligned and base models are close in distribution. Extensive experiments show that LAPD achieves an improvement 45.82% relative to the strongest existing baselines, yielding large and consistent gains across all settings.

文本检测对齐印记零样本统计推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。