arXiv:2507.02694cs.CL2025-07ACL被引 24

用新基准评估大模型能否发现科研论文的局限性

Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers

  • 构建限制类型分类体系,设计双数据集评测大模型识限能力
  • 引入文献检索增强,使模型反馈更基于已有研究
  • 适合希望提升论文评审效率的研究者和审稿人

同行评审是科学研究的核心,但随着论文数量激增,这一高成本过程面临巨大挑战。尽管大语言模型在多种科学任务中展现出潜力,其在识别论文局限性方面的应用仍缺乏系统研究。本文提出针对人工智能领域科研论文的局限性类型分类体系,并基于此构建首个全面的评测基准LimitGen,包含通过高质量论文受控扰动生成的合成数据集LimitGen-Syn,以及真实人类撰写的局限性集合LimitGen-Human。为提升模型识别能力,引入文献检索机制以确保判断基于已有科学发现。该方法显著增强了大模型生成具体、建设性反馈的能力,可为早期反馈和辅助人工审稿提供支持。

原文摘要 · Abstract (English)

Peer review is fundamental to scientific research, but the growing volume of publications has intensified the challenges of this expertise-intensive process. While LLMs show promise in various scientific tasks, their potential to assist with peer review, particularly in identifying paper limitations, remains understudied. We first present a comprehensive taxonomy of limitation types in scientific research, with a focus on AI. Guided by this taxonomy, for studying limitations, we present LimitGen, the first comprehensive benchmark for evaluating LLMs' capability to support early-stage feedback and complement human peer review. Our benchmark consists of two subsets: LimitGen-Syn, a synthetic dataset carefully created through controlled perturbations of high-quality papers, and LimitGen-Human, a collection of real human-written limitations. To improve the ability of LLM systems to identify limitations, we augment them with literature retrieval, which is essential for grounding identifying limitations in prior scientific findings. Our approach enhances the capabilities of LLM systems to generate limitations in research papers, enabling them to provide more concrete and constructive feedback.

大模型科研评审局限性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。