arXiv:2605.01647cs.CL2026-05

用字符分布差异检测AI文本,突破传统困惑度上限。

Beyond Perplexity: Character Distribution Signatures and the MDTA Benchmark for AI Text Detection

论文配图:Beyond Perplexity: Character Distribution Signatures and the MDTA Benchmark for AI Text Detection
图 1 · 摘自论文原文
  • 基于字符分布模式差异构建新检测信号,避开概率伪装
  • 在4个模型5个领域中验证,检测准确率显著提升
  • 适合研究生成文本真实性或对抗攻击的学者使用

无需训练的AI文本检测方法主要依赖模型的词元概率分布,在Binoculars和DNA-DetectLLM等方法上表现优异。然而,这些方法面临根本性瓶颈:大模型经RLHF优化后,其概率分布已趋近人类书写风格。本文提出一种基于字符分布签名的新检测信号。理论分析表明,训练于大规模领域均衡语料的AI模型会逼近全局字符模式,而人类则表现出领域特异性分布,形成“分离墙”——人机差异远大于机器间差异。为系统评估,我们构建了包含642,274个提示对齐样本的MDTA基准,涵盖4个模型、5个领域、3种温度设置和3种对抗策略,大幅扩展了HC3数据集。引入字母分布得分(LD-Score),其与困惑度方法相关性极低(r = 0.08–0.13)。将LD-Score与DNA-DetectLLM、Binoculars、FastDetectGPT结合,通过非线性分类器处理,持续提升AUROC与F1分数,尤其在词汇受限的专业领域效果更明显。数据集可访问:https://huggingface.co/datasets/nsp909/MDTA。

原文摘要 · Abstract (English)

Training-free AI text detection methods primarily rely on model log-probabilities, achieving strong performance through approaches like Binoculars and DNA-DetectLLM. However, these methods face a fundamental ceiling as models are optimized through RLHF to produce human-like probability distributions. We introduce an alternative detection signal based on character distribution signatures. We provide theoretical foundations showing that AI models, trained on massive domain-balanced corpora, approximate global character patterns while humans exhibit domain-specialized distributions, creating a "Wall of Separation" where human-AI divergence significantly exceeds AI-AI divergence. To enable systematic evaluation, we construct the Models-Domains-Temperatures-Adversarials (MDTA) benchmark comprising 642,274 prompt-aligned samples across 4 models, 5 domains, 3 temperature settings, and 3 adversarial strategies, substantially expanding the HC3 dataset with modern model responses, temperature variation, and adversarial augmentation. We introduce the Letter Distribution Score (LD-Score), demonstrating low correlation (r = 0.08-0.13) with perplexity methods. When integrated with DNA-DetectLLM, Binoculars and FastDetectGPT via a non-linear classifier, LD-Score yields consistent improvements in AUROC and F1, with particularly pronounced gains in specialized domains where vocabulary constraints amplify the detection signal. The MDTA dataset can be accessed at: https://huggingface.co/datasets/nsp909/MDTA.

文本检测字符分布鲁棒性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。