arXiv:2503.20421cs.CLcs.LG2025-03被引 3

通过检测生成文本的局部归一化缺陷,实现不依赖模型的机器文本识别。

TempTest: Local Normalization Distortion and the Detection of Machine-generated Text

  • 利用温度/Top-k采样导致的概率归一化偏差作为检测信号。
  • 在多种模型、数据集和文本长度下表现优于或媲美现有方法。
  • 对改写攻击鲁棒,且不对非母语者产生明显偏见。

现有的零样本机器生成文本检测方法主要依赖三个统计量:对数似然、对数排名和熵。随着语言模型越来越接近人类文本分布,这些方法的有效性将受限。为此,我们提出一种完全不依赖生成语言模型的检测方法,其核心是针对温度或Top-k采样等解码策略所引发的条件概率度量归一化缺陷。该方法具有严格的理论基础,易于解释,且在概念上区别于已有检测手段。我们在白盒与黑盒设置下,针对多种语言模型、数据集及文本长度进行了评估,并研究了改写攻击的影响以及对非母语者的潜在偏见。在所有测试场景中,本方法性能至少可与当前最优检测器比肩,部分情况下显著超越基线。

原文摘要 · Abstract (English)

Existing methods for the zero-shot detection of machine-generated text are dominated by three statistical quantities: log-likelihood, log-rank, and entropy. As language models mimic the distribution of human text ever closer, this will limit our ability to build effective detection algorithms. To combat this, we introduce a method for detecting machine-generated text that is entirely agnostic of the generating language model. This is achieved by targeting a defect in the way that decoding strategies, such as temperature or top-k sampling, normalize conditional probability measures. This method can be rigorously theoretically justified, is easily explainable, and is conceptually distinct from existing methods for detecting machine-generated text. We evaluate our detector in the white and black box settings across various language models, datasets, and passage lengths. We also study the effect of paraphrasing attacks on our detector and the extent to which it is biased against non-native speakers. In each of these settings, the performance of our test is at least comparable to that of other state-of-the-art text detectors, and in some cases, we strongly outperform these baselines.

文本检测生成模型归一化缺陷零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。