arXiv:2604.16607cs.CLcs.AI2026-04被引 2

评测15种文本检测模型,发现结果高度依赖数据集和评估指标。

Spotlights and Blindspots: Evaluating Machine-Generated Text Detection

  • 在7个文本集和3个创作类数据集上对比15种检测模型。
  • 无模型在所有场景下表现最优,且对新型人类写作文本检测效果差。
  • 评估方法选择显著影响结果,需谨慎设计实验流程。

随着生成式语言模型的兴起,机器生成文本检测成为关键挑战。尽管已有多种检测模型,但数据集不一致、评估指标多样及评估策略差异导致模型效果难以比较。为此,我们评估了来自六个不同系统的15种检测模型,以及七种训练好的模型,在七个英文文本测试集和三个创造性人类写作数据集上的表现。通过实证分析模型性能、训练与评估数据的影响,以及关键指标的作用,发现没有单一系统在所有任务中领先,几乎所有模型在特定任务中均有效;模型表现的呈现与数据集和指标选择密切相关。基于不同数据集和指标,模型排名变化显著,且在高风险领域的新颖人类写作文本上整体表现不佳。研究还表明,许多常被忽视的方法论选择对准确反映模型性能至关重要。

原文摘要 · Abstract (English)

With the rise of generative language models, machine-generated text detection has become a critical challenge. A wide variety of models is available, but inconsistent datasets, evaluation metrics, and assessment strategies obscure comparisons of model effectiveness. To address this, we evaluate 15 different detection models from six distinct systems, as well as seven trained models, across seven English-language textual test sets and three creative human-written datasets. We provide an empirical analysis of model performance, the influence of training and evaluation data, and the impact of key metrics. We find that no single system excels in all areas and nearly all are effective for certain tasks, and the representation of model performance is critically linked to dataset and metric choices. We find high variance in model ranks based on datasets and metrics, and overall poor performance on novel human-written texts in high-risk domains. Across datasets and metrics, we find that methodological choices that are often assumed or overlooked are essential for clearly and accurately reflecting model performance.

文本检测评估方法生成文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。