评测大模型生成论文评审的优劣,发现其长于描述优点,弱于指出问题。
Unveiling the Merits and Defects of LLMs in Automatic Review Generation for Scientific Papers
- 构建包含1683篇论文的基准数据集,用知识图谱分析模型生成评审的质量
- 大模型在好论文中多生成15.74%的正面内容实体,但弱论文中仅增5.7%反馈
- 所有模型均严重低估缺陷,仅59.42%的负面实体量,远低于人类50%差异
科学论文投稿量激增给传统同行评审带来压力,促使研究者探索大语言模型(LLMs)在自动化评审中的应用。尽管LLMs能生成结构清晰的反馈,但在批判性推理、上下文理解与质量敏感性方面仍有限。为此,我们提出一个综合评估框架,结合语义相似度分析与结构化知识图谱指标,对比模型与人工评审。构建了涵盖ICLR和NeurIPS多年共1683篇论文及6495份专家评审的大规模基准数据集,并使用五种LLMs生成评审。结果显示,LLMs在描述性与肯定性内容上表现良好,能准确捕捉论文主要贡献与方法;以GPT-4o为例,在ICLR 2025优质论文的亮点部分比人类多生成15.74%的实体。然而,在识别弱点、提出实质性问题以及根据论文质量调整反馈方面明显不足:其在弱点部分实体数量仅为人类的59.42%,且从优质到劣质论文的节点增长仅5.7%,而人类为50%。上述趋势在各会议、年份与模型间一致,为理解LLM评审的优缺点提供了实证基础,有助于未来辅助评审工具的改进。数据、代码与详细结果公开于https://github.com/RichardLRC/Peer-Review。
原文摘要 · Abstract (English)
The surge in scientific submissions has placed increasing strain on the traditional peer-review process, prompting the exploration of large language models (LLMs) for automated review generation. While LLMs demonstrate competence in producing structured and coherent feedback, their capacity for critical reasoning, contextual grounding, and quality sensitivity remains limited. To systematically evaluate these aspects, we propose a comprehensive evaluation framework that integrates semantic similarity analysis and structured knowledge graph metrics to assess LLM-generated reviews against human-written counterparts. We construct a large-scale benchmark of 1,683 papers and 6,495 expert reviews from ICLR and NeurIPS in multiple years, and generate reviews using five LLMs. Our findings show that LLMs perform well in descriptive and affirmational content, capturing the main contributions and methodologies of the original work, with GPT-4o highlighted as an illustrative example, generating 15.74% more entities than human reviewers in the strengths section of good papers in ICLR 2025. However, they consistently underperform in identifying weaknesses, raising substantive questions, and adjusting feedback based on paper quality. GPT-4o produces 59.42% fewer entities than real reviewers in the weaknesses and increases node count by only 5.7% from good to weak papers, compared to 50% in human reviews. Similar trends are observed across all conferences, years, and models, providing empirical foundations for understanding the merits and defects of LLM-generated reviews and informing the development of future LLM-assisted reviewing tools. Data, code, and more detailed results are publicly available at https://github.com/RichardLRC/Peer-Review.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。