arXiv:2511.11340cs.CLcs.AI2025-11被引 10

构建多领域AI生成文本检测任务,推动真假内容识别研究

M-DAIGT: A Shared Task on Multi-Domain Detection of AI-Generated Text

  • 设立新闻与学术文本两类检测子任务,评估跨领域识别能力
  • 发布3万样本平衡数据集,涵盖多种大模型与提示策略生成内容
  • 4支队伍参与,验证主流检测方法在真实场景的适用性

大型语言模型生成高度流畅文本,对信息真实性和学术研究构成严峻挑战。本文提出多领域AI生成文本检测共享任务(M-DAIGT),聚焦新闻文章与学术写作中的AI生成文本识别。该任务包含两个二分类子任务:新闻文章检测(NAD,子任务1)和学术写作检测(AWD,子任务2)。为支持此任务,我们构建并发布了包含3万条样本的大规模基准数据集,人类撰写与AI生成文本比例均衡。AI内容由多种现代大模型(如GPT-4、Claude)及多样化提示策略生成。共有46支团队注册,其中4支提交最终结果,全部参与两个子任务。本文介绍参赛团队所采用的方法,并简要讨论M-DAIGT的未来方向。

原文摘要 · Abstract (English)

The generation of highly fluent text by Large Language Models (LLMs) poses a significant challenge to information integrity and academic research. In this paper, we introduce the Multi-Domain Detection of AI-Generated Text (M-DAIGT) shared task, which focuses on detecting AI-generated text across multiple domains, particularly in news articles and academic writing. M-DAIGT comprises two binary classification subtasks: News Article Detection (NAD) (Subtask 1) and Academic Writing Detection (AWD) (Subtask 2). To support this task, we developed and released a new large-scale benchmark dataset of 30,000 samples, balanced between human-written and AI-generated texts. The AI-generated content was produced using a variety of modern LLMs (e.g., GPT-4, Claude) and diverse prompting strategies. A total of 46 unique teams registered for the shared task, of which four teams submitted final results. All four teams participated in both Subtask 1 and Subtask 2. We describe the methods employed by these participating teams and briefly discuss future directions for M-DAIGT.

文本检测AI生成大模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。