arXiv:2606.15887cs.LGcs.AI2026-06综述

用大模型评分论文,效果比人工评审还准。

Intelligence Is Not the Bottleneck: Validating an LLM First-Pass Manuscript Score Against Peer-Review Outcomes

  • 仅靠提示词生成论文评分,不微调直接预测。
  • 评分能准确区分录用与拒稿,AUROC达0.82。
  • 适合想快速评估论文质量的研究者使用。

大型语言模型(LLM)系统被广泛用于辅助同行评审,但多数评估关注机器生成评审文本的文笔,而非其评分的有效性。本文验证AIPR系统——该系统读取投稿论文后输出五个0-100分的质量维度及加权总分——在顶级机器学习会议ICLR的公开决策结果上的表现。基于300篇有公开决策层级和评审评分的ICLR投稿,采用冻结流程、事前注册假设,在任何评分与结果关联前完成实验。总分能有效区分拒稿与录用(AUROC 0.82,95%置信区间0.78–0.87),随层级单调上升,并与平均评审评分一致。得分最低的五分之一稿件被拒率远高于基线,且无口头报告论文。有效性主要来自模型本身:一段提示词即可实现接近完整流程的判别能力(小差距未达预设显著性标准,p=0.09)。工程部分提升的是稳定性和可解释性:同一论文重复运行时,评分波动极小(组内标准差0.7 vs. 裸提示2.8),并生成结构化、有依据的评审意见,而人类仍掌握最终决策权。

原文摘要 · Abstract (English)

Large language model (LLM) systems are increasingly proposed to assist peer review, yet most evaluations judge the prose of machine-generated review text, not the validity of the numeric score a system assigns. We validate AIPR, which reads a submitted manuscript and emits five 0-100 quality dimensions and a weighted overall score, against the public decision outcomes of a major machine learning venue. AIPR grades by prompting alone, with no fine-tuning on reviews or decisions. Across 300 ICLR submissions with public decision tiers and reviewer ratings, graded under a frozen pipeline with hypotheses pre-registered before any score met any outcome, the overall score separates rejected from accepted submissions (AUROC 0.82, 95% CI 0.78-0.87), rises monotonically across tiers, and tracks the mean reviewer rating. The signal is strongest where we claim it: the lowest-scoring fifth is rejected far above the base rate, with oral papers absent. The validity comes mostly from the model: a one-paragraph prompt on the same model discriminates almost as well as the full pipeline (the small gap favours the pipeline but does not meet the pre-declared criterion, p = 0.09). What the engineering adds is reliability and a grounded review: AIPR's score barely moves across repeated runs (0.7 vs. 2.8 points within-paper SD) where the bare prompt swings, and the same pass returns a rubric-structured, evidence-grounded review rather than a bare number, with the human keeping the decision.

大模型评分同行评审论文评估AI辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。