用大模型检查论文结论是否被方法支撑,提升审稿准确性。
Do Methods Support the Claims? Intra-Paper Verification for Peer Review

- 通过提取论文创新点与方法证据,判断两者是否匹配。
- 在ICLR 2025论文上验证,与人类审稿人意见高度一致。
- 适合用于辅助审稿、提升论文可信度评估效率。
科学论文数量激增促使使用大语言模型(LLM)辅助同行评审。现有自动化新颖性评估方法通常将论文声称的贡献与已有文献对比,隐含假设这些贡献在论文中已被准确实现。然而,人类审稿人常质疑新颖性,并非因已有相似想法,而是因为文中方法论证据不足以支撑其主张。这种主张与方法实现之间的内部不一致,当前基于LLM的评审系统极少关注。为此,我们提出“论文内主张验证”框架,评估论文中宣称的创新是否由实际采用的方法充分支持。该框架利用LLM从引言中提取新颖性主张,检索相关方法证据,并判断方法能否支撑所述贡献。评估依据182篇ICLR 2025论文收集的人类审稿意见归纳出的评价标准,涵盖新颖性、方法、清晰度等常见问题,生成结构化审稿式评估。通过对比框架生成评论与人工审稿人关注点,在已接受和拒绝论文的平衡样本上进行评估,结果显示框架与人类审稿人意见高度吻合,尤其在新颖性问题上表现突出。BERTScore进一步区分了对应的人工-模型评论对与不匹配对照组,表明该框架捕捉到了与人类观察一致的关切点。
原文摘要 · Abstract (English)
The growing volume of scientific submissions has motivated interest in using large language models (LLMs) to assist peer review. Existing automated novelty assessment approaches typically compare a paper's claimed contributions against prior literature, implicitly assuming that these contributions are accurately realized in the work itself. Human reviewers, however, frequently challenge novelty claims not because similar ideas already exist, but because the methodological evidence presented in the paper does not adequately support them. This internal mismatch between claimed contributions and methodological realization is rarely examined by current LLM-based review systems. To address this gap, we introduce intra-paper claim verification, a framework that evaluates whether novelty claims articulated in a paper are substantiated by the methods used to realize them. The framework employs an LLM to extract novelty claims from the introduction, retrieve claim-relevant methodological evidence, and assess whether the methods substantiate the stated contributions. Assessment is guided by reviewer-inspired evaluation criteria derived inductively from human peer reviews collected from 182 ICLR 2025 papers. These criteria capture recurring reviewer concerns related to novelty, methodology, clarity, and other issues and are used to generate structured reviewer-style assessments of claim substantiation. We evaluate the framework by comparing LLM-generated review comments against human reviewer concerns on a balanced subset of accepted and rejected papers. Human evaluation demonstrates significant alignment between framework-generated assessments and human reviewer concerns, particularly for novelty-related issues. BERTScore further distinguishes corresponding human-LLM review pairs from mismatched controls, indicating that the framework captures concerns consistent with human reviewer observations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。