用细粒度关切分析取代简单对错判断,提升AI评审可信度。
What Makes a Good AI Review? Concern-Level Diagnostics for AI Peer Review

- 构建匹配图,标注关切类型、严重性及回应后处理,实现关切级诊断
- 发现多数系统将25%~55%已接收论文的关切标记为决定性,与实际不符
- 强调优先级校准比检测率更重要,适合评审系统开发者和评测者
评估AI生成评审意见的结论一致性被广泛认为不足,而现有替代方案很少审计系统识别了哪些关切、如何排序,以及这些排序是否与最终评估理由一致。本文提出关切对齐(concern alignment)诊断框架,从关切层面而非仅结论层面评估AI评审。核心数据结构是匹配图,即官方与AI生成关切间的二分图对齐,包含匹配类型、严重性和回应后处理信息。由此衍生出评估阶梯:从二元准确率到关切检测、结论分层行为、决策感知校准,再到回应感知分解。在四款公开评审系统六种配置的试点研究中,关切级分析表明,仅检测关切不足以决定评审质量;校准常为关键瓶颈。系统虽检测到非零比例的官方关切,但多数将25%–55%已接收论文的关切标为决定性,而根据操作定义,接收论文无官方关切被视作决定性障碍。相同总体结论准确率可能掩盖拒绝偏好与低召回模式的差异,且低全评误判率可能源于关切稀释而非精准校准。多数系统不原生输出接受/拒绝,通过语气推断易受方法影响,凸显关切级诊断在不同推断方式下仍稳定的必要性。贡献在于提供可复用的评估框架,用于审计AI评审系统识别了哪些关切、如何加权,以及加权是否与评审理由一致。
原文摘要 · Abstract (English)
Evaluating AI-generated reviews by verdict agreement is widely recognized as insufficient, yet current alternatives rarely audit which concerns a system identifies, how it prioritizes them, or whether those priorities align with the review rationale that shaped the final assessment. We propose concern alignment, a diagnostic framework that evaluates AI reviews at the concern level rather than only at the verdict level. The framework's core data structure is the match graph, a bipartite alignment between official and AI-generated concerns annotated with match type, severity, and post-rebuttal treatment. From this artifact we derive an evaluation ladder that moves from binary accuracy to concern detection, verdict-stratified behavior, decision-aware calibration, and rebuttal-aware decomposition. In a pilot study of four public AI review systems evaluated in six configurations, concern-level analysis suggests that detection alone does not determine review quality; calibration is often the binding constraint. Systems detect non-trivial fractions of official concerns yet most mark 25--55% of concerns on accepted papers as decisive, where, under our operationalization, no official concern on accepted papers was treated as a decisive blocker. Identical overall verdict accuracy can conceal reject-heavy behavior versus low-recall profiles, and low full-review false decisive rates can partly reflect concern dilution rather than calibrated prioritization. Most systems do not emit a native accept/reject, and inferring it from review tone is method-sensitive, reinforcing the need for concern-level diagnostics that remain stable across inference choices. The contribution is a reusable evaluation framework for auditing which concerns AI reviewers identify, how they weight them, and whether those priorities align with the review rationale that informed the paper's final assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。