arXiv:2505.23824cs.CL2025-05中稿 · NeurIPS被引 15

用大模型自动检查论文关键错误,提升审稿效率与可靠性

Reviewing Scientific Papers for Critical Problems With Reasoning LLMs: Baseline Approaches and Automatic Evaluation

  • 用推理型大模型充当论文质量检查员,避免生成完整评审
  • o3模型在识别关键错误上表现最佳,成本适中
  • 提供可复现的评估框架,适合研究可信AI审稿者的人

大型语言模型的进展激发了利用其辅助科学出版同行评审的热潮,但直接由AI生成完整评审可能加剧不负责任的使用并引发故意操纵。为此,我们提出将大模型作为手稿质量检查工具。本文介绍多种基线方法和一个可扩展的自动评估框架,采用顶尖推理型大模型作为评判者,解决领域专家难以招募的问题。基于从arXiv撤稿的论文,我们在2025年5月至6月期间验证了所提方法,评估了多个领先推理型大模型在识别科学论文中的关键错误与逻辑缺陷方面的表现及API成本。结果显示,o3模型在问题识别性能上优于所有其他模型,且成本可控。本研究为文档级科学理解与推理提供了洞见,并为未来应用奠定基础。数据集、代码与模型输出均已公开。

原文摘要 · Abstract (English)

Recent advancements in large language models have sparked interest in utilizing them to aid the peer review process of scientific publication amid the peer review crisis. However, having AI models generate full reviews in the same way as human reviewers risks exacerbating the irresponsible use of LLM-generated reviews and instigating intentional manipulation. As an alternative, we propose adopting LLMs as manuscript quality checkers. We introduce several baseline approaches and an extendable automatic evaluation framework using top reasoning LLMs as judges to tackle the difficulty of recruiting domain experts for manual evaluation. Utilizing papers withdrawn from arXiv, we validated our proposed methods with several leading reasoning LLMs available in May-June 2025 and assessed their performance and API costs for identifying critical errors and unsoundness problems in scientific papers. o3 exhibited the best problem identification performance among all models at a modest cost. This paper provides insights into document-based scientific understanding/reasoning and lays a foundation for future applications. Our dataset, code, and model outputs are publicly available.

AI审稿大模型评估科学写作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。