用五维评估框架提升AI评阅质量,强调文本论证比评分更重要
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews

- 构建五维评估体系,从论证、聚焦、提问等角度衡量AI评阅质量
- 实验显示,弱点评述召回率与评分准确率高度相关
- 适合关注AI审稿可靠性、需改进评阅机制的研究者
大型语言模型的快速应用推动了自动化同行评审的发展,但当前评估基准将评审简化为评分预测任务,限制了进展。我们主张,评审的价值在于其文本论据——即观点、问题与批判性分析,而非单一分数。为此,我们提出Beyond Rating框架,从五个维度评估AI评审:内容忠实度、论证一致性、焦点稳定性、提问建设性及人类倾向性。特别地,我们采用最大召回策略应对专家间合理分歧,并构建了一个经严格筛选的高质量论文-高置信度评审数据集,剔除流程噪声。大量实验证明,传统n-gram指标无法反映人类偏好,而以弱点评述召回率为代表的文本中心指标与评分准确性高度相关。结果表明,使AI评阅焦点与人类专家一致是实现可靠自动评分的前提,为未来研究提供了可靠标准。
原文摘要 · Abstract (English)
The rapid adoption of Large Language Models (LLMs) has spurred interest in automated peer review; however, progress is currently stifled by benchmarks that treat reviewing primarily as a rating prediction task. We argue that the utility of a review lies in its textual justification--its arguments, questions, and critique--rather than a scalar score. To address this, we introduce Beyond Rating, a holistic evaluation framework that assesses AI reviewers across five dimensions: Content Faithfulness, Argumentative Alignment, Focus Consistency, Question Constructiveness, and AI-Likelihood. Notably, we propose a Max-Recall strategy to accommodate valid expert disagreement and introduce a curated dataset of paper with high-confidence reviews, rigorously filtered to remove procedural noise. Extensive experiments demonstrate that while traditional n-gram metrics fail to reflect human preferences, our proposed text-centric metrics--particularly the recall of weakness arguments--correlate strongly with rating accuracy. These findings establish that aligning AI critique focus with human experts is a prerequisite for reliable automated scoring, offering a robust standard for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。