分析论文评审中的细粒度矛盾,判断冲突强度与证据。
When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews

- 以整篇评审为单位,识别矛盾证据并打分冲突强度。
- 提出新数据集RevCI和模型IMPACT,性能显著优于基线。
- 轻量模型TIDE可单次推理完成,适合实际部署使用。
科学同行评审常出现专家意见冲突,随着会议投稿量增加,领域主席和编辑难以可靠识别和解读此类分歧。现有方法通常将评审分歧简化为孤立句子对的二元矛盾检测,忽略了评审层面的上下文,也掩盖了评价冲突的严重程度差异。本文提出一种细粒度评审矛盾分析范式,基于完整评审文本显式识别矛盾证据片段,并分配分级的分歧强度评分。为此,我们构建了RevCI——一个由专家标注的评审对数据集,包含证据级矛盾标注和分级强度标签。进一步提出IMPACT框架,采用结构化多智能体设计,融合方面条件的证据提取、思辨推理与裁决机制,有效建模评审矛盾及其强度。为支持高效部署,我们将IMPACT蒸馏为小模型TIDE,可在单次前向传播中预测矛盾证据与强度。实验表明,IMPACT在证据识别和强度一致率上显著超越强基线模型;而TIDE在推理成本大幅降低的同时仍保持竞争力。
原文摘要 · Abstract (English)
Scientific peer reviews frequently contain conflicting expert judgments, and the increasing scale of conference submissions makes it challenging for Area Chairs and editors to reliably identify and interpret such disagreements. Existing approaches typically frame reviewer disagreement as binary contradiction detection over isolated sentence pairs, abstracting away the review-level context and obscuring differences in the severity of evaluative conflict. In this work, we introduce a fine-grained formulation of reviewer contradiction analysis that operates over full peer reviews by explicitly identifying contradiction evidence spans and assigning graded disagreement intensity scores. To support this task, we present RevCI, an expert-annotated benchmark of peer-review pairs with evidence-level contradiction annotations with graded intensity labels. We further propose IMPACT, a structured multi-agent framework that integrates aspect-conditioned evidence extraction, deliberative reasoning, and adjudication to model reviewer contradictions and their intensity. To support efficient deployment, we distill IMPACT into TIDE, a small language model that predicts contradiction evidence and intensity in a single forward pass. Experimental results show that IMPACT substantially outperforms strong single-agent and generic multi-agent baselines in both evidence identification and intensity agreement, while TIDE achieves competitive performance at significantly lower inference cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。