构建俄语科学摘要的AI生成检测数据集,推动多领域跨模型识别研究。
AINL-Eval 2025 Shared Task: Detection of AI-Generated Scientific Abstracts in Russian
- 构建52,305样本的俄语科学摘要数据集,含12领域真人与5大模型生成内容。
- 挑战模型在未见领域和未训练模型上的泛化能力,顶尖系统表现优异。
- 适合关注AI内容检测、多语言NLP及学术诚信的研究者使用。
大型语言模型(LLMs)的快速发展使文本生成日益逼真,难以区分人类与AI生成内容,对学术诚信构成严峻挑战,尤其在多语言环境中检测资源匮乏。为此,我们推出AINL-Eval 2025共享任务,聚焦俄语科学摘要的AI生成检测。构建了一个包含52,305个样本的大规模数据集,涵盖12个不同科学领域的人类撰写摘要及来自五种先进LLM(GPT-4-Turbo、Gemma2-27B、Llama3.3-70B、Deepseek-V3和GigaChat-Lite)的生成对应物。任务核心目标是要求参赛者开发具备泛化能力的解决方案,能应对(i)此前未见过的科学领域和(ii)训练数据中未包含的模型。任务分为两阶段,吸引10支团队提交159份方案,最优系统在识别AI生成内容上表现突出。我们还建立了持续运行的共享任务平台,以促进该领域长期研究进展。数据集与平台已公开于https://github.com/iis-research-team/AINL-Eval-2025。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) has revolutionized text generation, making it increasingly difficult to distinguish between human- and AI-generated content. This poses a significant challenge to academic integrity, particularly in scientific publishing and multilingual contexts where detection resources are often limited. To address this critical gap, we introduce the AINL-Eval 2025 Shared Task, specifically focused on the detection of AI-generated scientific abstracts in Russian. We present a novel, large-scale dataset comprising 52,305 samples, including human-written abstracts across 12 diverse scientific domains and AI-generated counterparts from five state-of-the-art LLMs (GPT-4-Turbo, Gemma2-27B, Llama3.3-70B, Deepseek-V3, and GigaChat-Lite). A core objective of the task is to challenge participants to develop robust solutions capable of generalizing to both (i) previously unseen scientific domains and (ii) models not included in the training data. The task was organized in two phases, attracting 10 teams and 159 submissions, with top systems demonstrating strong performance in identifying AI-generated content. We also establish a continuous shared task platform to foster ongoing research and long-term progress in this important area. The dataset and platform are publicly available at https://github.com/iis-research-team/AINL-Eval-2025.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。