arXiv:2512.05925cs.AIcs.CL2025-12被引 5

用大模型检测顶会论文中的客观错误,发现错误率逐年上升。

To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis

  • 用GPT-5自动分析顶会论文中的公式、图表等客观错误
  • 2021到2025年论文平均错误数从3.8升至5.9,增幅55.3%
  • 大模型可修复75.8%的错误,适合关注可复现性的研究者

发表的AI论文中存在多少错误?同行评审文献是新研究的基础,但其中的错误若未被发现,会引发后续研究混淆并影响可复现性。为应对这一问题,我们基于GPT-5开发了论文正确性检查工具,系统分析顶级会议和期刊已发表论文中的客观错误——如公式、推导、计算、图表和表格中的错误,这些错误具有明确可验证的真值。我们排除了新颖性、重要性等主观判断。结果发现,论文中存在显著数量的客观错误,且平均错误数呈上升趋势:NeurIPS 2021为3.8,2025年升至5.9(增加55.3%);ICLR 2018为4.1,2025年为5.2;TMLR 2022/23为5.0,2025年为5.5。人工专家审核316个疑似错误,确认263个为真实错误,准确率达83.2%。多数问题较轻微,但纠正后可减少文献混乱、提升可复现性。检查工具还发现了可能影响结果解释的更严重错误,并能对75.8%的错误提出正确修正方案。本研究证明前沿大模型在识别与修正论文客观错误方面的潜力,有助于建立更稳固的知识基础。

原文摘要 · Abstract (English)

How many mistakes do published AI papers contain? Peer-reviewed publications form the foundation upon which new research and knowledge are built. Errors that persist in the literature can propagate unnoticed, creating confusion in follow-up studies and complicating reproducibility. The accelerating pace of research and the increasing demands on the peer-review system make such mistakes harder to detect and avoid. To address this, we developed a Paper Correctness Checker based on GPT-5 to systematically identify mistakes in papers previously published at top AI conferences and journals. Our analysis focuses on objective mistakes-e.g., errors in formulas, derivations, calculations, figures, and tables-that have a clearly verifiable ground truth. We intentionally exclude subjective considerations such as novelty, importance, or writing quality. We find that published papers contain a non-negligible number of objective mistakes and that the average number of mistakes per paper has increased over time-from 3.8 in NeurIPS 2021 to 5.9 in NeurIPS 2025 (55.3% increase); from 4.1 in ICLR 2018 to 5.2 in ICLR 2025; and from 5.0 in TMLR 2022/23 to 5.5 in TMLR 2025. Human experts reviewed 316 potential mistakes identified by the AI Checker and confirmed that 263 were actual mistakes, corresponding to a precision of 83.2%. While most identified issues are relatively minor, correcting them would reduce confusion in the literature and strengthen reproducibility. The AI Checker also surfaced potentially more substantive mistakes that could affect the interpretation of results. Moreover, we show that the AI Checker can propose correct fixes for 75.8% of the identified mistakes. Overall, this study highlights the potential of frontier LLMs to detect and correct objective mistakes in published papers, helping to establish a firmer foundation of knowledge.

大模型评估可复现性错误检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。