arXiv:2510.22874cs.CL2025-10被引 3

构建超7万条真实与AI生成文本对比数据集,助力检测与溯源。

A Comprehensive Dataset for Human vs. AI Generated Text Detection

  • 融合《纽约时报》原文与多款主流大模型生成文本
  • 实现58.35%的真假文本识别准确率,模型溯源准确率达8.92%
  • 适合研究生成内容检测、AI可信度评估的学者与工程师

大型语言模型的快速发展使得AI生成文本愈发接近人类写作,引发对内容真实性、虚假信息和可信度的担忧。可靠检测AI生成文本并追溯其来源,亟需大规模、多样化且标注清晰的数据集。本文构建了一个包含超过73,193个文本样本的综合性数据集,将《纽约时报》的真实文章与Gemma-2-9b、Mistral-7B、Qwen-2-72B、LLaMA-8B、Yi-Large及GPT-4-o等多款先进大模型生成的合成文本相结合。数据集以文章摘要为提示,提供完整的真人撰写文本。我们建立了两个关键任务的基线:区分真人与AI文本的准确率为58.35%,识别生成模型的准确率为8.92%。该数据集通过连接真实新闻内容与现代生成模型,旨在推动鲁棒检测与溯源方法的发展,提升生成式AI时代下的信任与透明度。数据集已公开于:https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Text_Dataset

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authenticity, misinformation, and trustworthiness. Addressing the challenge of reliably detecting AI-generated text and attributing it to specific models requires large-scale, diverse, and well-annotated datasets. In this work, we present a comprehensive dataset comprising over 73,193 text samples that combine authentic New York Times articles with synthetic versions generated by multiple state-of-the-art LLMs including Gemma-2-9b, Mistral-7B, Qwen-2-72B, LLaMA-8B, Yi-Large, and GPT-4-o. The dataset provides original article abstracts as prompts, full human-authored narratives. We establish baseline results for two key tasks: distinguishing human-written from AI-generated text, achieving an accuracy of 58.35\%, and attributing AI texts to their generating models with an accuracy of 8.92\%. By bridging real-world journalistic content with modern generative models, the dataset aims to catalyze the development of robust detection and attribution methods, fostering trust and transparency in the era of generative AI. Our dataset is available at: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Text_Dataset

文本检测AI溯源数据集大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。