arXiv:2412.18274cs.CLcs.AI2024-12被引 20

检测学术作文是人工写的还是AI生成的,英语和阿拉伯语都可用。

GenAI Content Detection Task 2: AI vs. Human -- Academic Essay Authenticity Challenge

  • 用微调的Transformer模型或LLM判断作文作者
  • 最佳系统在中英文上F1超0.98,远超基础方法
  • 适合关注AI内容检测的研究者和教育从业者

本文全面介绍首届学术作文真实性挑战赛,该赛事作为COLING 2025期间的GenAI内容检测共用任务之一。挑战赛聚焦于识别学术用途下的机器生成与人工撰写作文,任务定义为:给定一篇作文,判断其是否由机器生成。比赛涵盖英语和阿拉伯语两种语言。评估阶段共有25支团队提交英语系统,21支团队提交阿拉伯语系统,反映出高度关注。最终有七支队伍提交了系统说明论文。多数参赛系统采用微调的Transformer模型,其中一支团队使用Llama 2和Llama 3等大型语言模型。本文详述任务设定、数据集构建过程及评估框架,并总结各队采用的方法。几乎所有系统均优于基于n-gram的基线模型,顶尖系统在两种语言上的F1分数均超过0.98,表明机器生成文本检测已取得显著进展。

原文摘要 · Abstract (English)

This paper presents a comprehensive overview of the first edition of the Academic Essay Authenticity Challenge, organized as part of the GenAI Content Detection shared tasks collocated with COLING 2025. This challenge focuses on detecting machine-generated vs. human-authored essays for academic purposes. The task is defined as follows: "Given an essay, identify whether it is generated by a machine or authored by a human.'' The challenge involves two languages: English and Arabic. During the evaluation phase, 25 teams submitted systems for English and 21 teams for Arabic, reflecting substantial interest in the task. Finally, seven teams submitted system description papers. The majority of submissions utilized fine-tuned transformer-based models, with one team employing Large Language Models (LLMs) such as Llama 2 and Llama 3. This paper outlines the task formulation, details the dataset construction process, and explains the evaluation framework. Additionally, we present a summary of the approaches adopted by participating teams. Nearly all submitted systems outperformed the n-gram-based baseline, with the top-performing systems achieving F1 scores exceeding 0.98 for both languages, indicating significant progress in the detection of machine-generated text.

AI检测学术写作多语言评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。