arXiv:2505.11855cs.CL2025-05被引 24

测试AI验证论文错误能力,发现现有模型效果极差。

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research

  • 构建包含91个重大错误的SPOT数据集,用于评估AI验证论文的能力。
  • 顶级模型最高仅21.1%召回率,多数接近零,且结果不稳定。
  • 错误类型类似学生误解,不适合用于学术严谨性审查。

大型语言模型(LLMs)正被视作科学发现的合作者,但现有研究多聚焦于生成假设或撰写论文。本文探索其作为验证者的潜力:自动核查科学论文的学术正确性。为此提出SPOT数据集,包含83篇已发表论文及其91个足以引发勘误或撤稿的重大错误,经作者与人工标注者交叉验证。在该数据集上评估主流LLM,发现无一模型召回率超过21.1%(o3表现最佳),精度最高仅6.1%,其余基本为零。模型置信度普遍偏低,八次独立运行中极少重复发现相同错误,可靠性堪忧。专家定性分析显示,即使最强模型也犯下类似学生水平的误解,源于对概念的误读。这些结果揭示当前LLM在可靠学术验证方面存在巨大差距。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have fueled the vision of automated scientific discovery, often called AI Co-Scientists. To date, prior work casts these systems as generative co-authors responsible for crafting hypotheses, synthesizing code, or drafting manuscripts. In this work, we explore a complementary application: using LLMs as verifiers to automate the \textbf{academic verification of scientific manuscripts}. To that end, we introduce SPOT, a dataset of 83 published papers paired with 91 errors significant enough to prompt errata or retraction, cross-validated with actual authors and human annotators. Evaluating state-of-the-art LLMs on SPOT, we find that none surpasses 21.1\% recall or 6.1\% precision (o3 achieves the best scores, with all others near zero). Furthermore, confidence estimates are uniformly low, and across eight independent runs, models rarely rediscover the same errors, undermining their reliability. Finally, qualitative analysis with domain experts reveals that even the strongest models make mistakes resembling student-level misconceptions derived from misunderstandings. These findings highlight the substantial gap between current LLM capabilities and the requirements for dependable AI-assisted academic verification.

AI验证论文纠错大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。