测试发现自动审稿模型无法识别论文中的逻辑错误。
Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
- 构建反事实评估框架,精准测试自动审稿模型的逻辑判断能力。
- 实验证明,论文逻辑缺陷对自动审稿输出无显著影响。
- 开源数据集与框架,助力提升审稿模型可靠性研究。
大型语言模型(LLMs)在加速和辅助学术同行评审方面潜力巨大,正被广泛用作全自动审稿生成器(ARGs)。然而,潜在偏见和系统性错误可能严重威胁科学诚信;理解当前先进ARGs的能力与局限至关重要。本文聚焦高质量同行评审的核心能力——识别研究逻辑错误,即评估论文结果、解释与主张之间的内部一致性。我们提出一个完全自动化的反事实评估框架,在受控条件下隔离并测试该能力。对多种ARG方法的测试显示,出人意料的是,研究逻辑缺陷对其输出审稿无显著影响。基于此,我们提出三项可操作建议,并公开发布反事实数据集与评估框架。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have great potential to accelerate and support scholarly peer review and are increasingly used as fully automatic review generators (ARGs). However, potential biases and systematic errors may pose significant risks to scientific integrity; understanding the specific capabilities and limitations of state-of-the-art ARGs is essential. We focus on a core reviewing skill that underpins high-quality peer review: detecting faulty research logic. This involves evaluating the internal consistency between a paper's results, interpretations, and claims. We present a fully automated counterfactual evaluation framework that isolates and tests this skill under controlled conditions. Testing a range of ARG approaches, we find that, contrary to expectation, flaws in research logic have no significant effect on their output reviews. Based on our findings, we derive three actionable recommendations for future work and release our counterfactual dataset and evaluation framework publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。