新基准评估AI评审过程可靠性,发现生成建议常不支持最终决定。
Beyond Final Decisions: A Process-Centric Benchmark for Transparent AI-Assisted Peer Review

- 构建基于评审流程的诊断基准,追踪从论文到决策的完整链条。
- 实验显示人工标注流程的决策价值显著高于模型生成结果。
- 适合关注AI评审可解释性与可信度的研究者使用。
同行评审是科学质量控制的核心。现有AI辅助评审评估多聚焦于评审质量或最终决策准确性,难以判断模型决策是否基于充分可靠的评审证据。本文提出一个以过程为中心的诊断基准,采用 (x, z_s, z_c, z_r, y) 表示论文内容、摘要、批判、建议与决策。将来自 PeerRead、NLPeer ARR-22 与 OpenReview-ICLR 的异构评审记录转换为流程对齐数据。以直接从论文内容预测决策(Direct)为基线,比较人工黄金流程(Gold-process)与模型预测流程(Predicted-process)的决策价值,并开展阶段级评估、链路一致性分析与干预敏感性测试。在三个数据集和六种模型上实验表明,黄金流程变量的决策价值普遍更高。主分析模型中,金标-预测差距在不同数据集和随机种子下保持稳定,且多数模型-数据组合中均重现该差距。尽管模型生成的中间文本在相邻阶段表现出较高局部一致性,但最终决策却未被前序评审证据一致支持。该基准面向辅助而非替代人类评审的AI系统,提供透明可审计的诊断工具,用于评估其评审过程的可靠性。
原文摘要 · Abstract (English)
Peer review is central to quality control in science. However, existing evaluations of AI-assisted peer review mainly focus on the overall quality of generated reviews or the accuracy of final decisions. They therefore provide limited evidence about whether model decisions are supported by sufficient and reliable review evidence. We introduce a process-centric diagnostic benchmark for AI-assisted peer review. It uses (x,$z_s$,$z_c$,$z_r$,y) to represent the paper content, summary, critique, suggestion, and decision. We convert heterogeneous review records from PeerRead, NLPeer ARR-22, and OpenReview-ICLR into process-aligned data. Our benchmark uses direct decision prediction from the paper content (Direct) as its baseline. It compares the decision value of Gold-process variables and Predicted-process variables, and conducts stage-level evaluation, chain-consistency evaluation, and interventional sensitivity analysis. Experiments across three datasets and six models show that Gold-process variables generally have higher decision value. For the main analysis model, the Gold--Predicted gap remains stable across datasets and random seeds. This gap is also reproduced in most model--dataset combinations. Although model-generated intermediate review texts show relatively high local consistency across adjacent stages, the final decisions are not consistently supported by the preceding review evidence. Our benchmark targets AI systems designed to assist rather than replace human reviewers. It provides a transparent and auditable diagnostic tool for evaluating the reliability of their review processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。