评估AI写审稿意见的质量,发现其存在偏乐观、泛化批评等缺陷。
AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

- 对比111个顶会与医学期刊的AI审稿政策,发现两领域监管差异大。
- 在ICLR 2026和Nature Communications上测试,发现模型生成审稿意见流畅但证据不足。
- 提出多维评估框架,避免仅靠总分误判AI审稿质量。
AI辅助同行评审正日益被讨论和采纳,作为支持科学出版流程的工具,但目前对出版平台如何规范其使用,以及当前AI评审系统的能力仍缺乏系统理解。本文首先调研了111个领先的人工智能/NLP会议与医学期刊的面向审稿人的AI政策,揭示了两个领域间显著的监管差异。其次,基于包含原始投稿稿件及数百份人工与机器生成审稿意见的新数据集,我们在ICLR 2026和Nature Communications上评估了AI生成的同行评审意见。采用互补评估指标(包括LLM-as-a-Judge、评分一致性、细节粒度、与人类审稿人关注点的重叠度),比较开源与专有模型的表现。结果表明,当前大语言模型可生成详尽且流畅的评审意见,但存在系统性弱点:推荐倾向过于积极、批评内容泛化、证据支撑不均。研究证明,仅依赖聚合质量分数可能高估评审质量,强调需进行多维度评估。
原文摘要 · Abstract (English)
AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are. We address these questions by first surveying reviewer-facing AI policies across 111 leading AI/NLP conferences and medical journals, revealing substantial regulation differences between the two communities. Second, we evaluate AI-generated peer reviews at ICLR 2026 and Nature Communications using a novel dataset comprising original manuscript submissions and several hundred human- and machine-generated reviews. We compare reviews produced by open-source and proprietary models using complementary evaluation metrics, including LLM-as-a-Judge, score alignment, granularity, and overlap with human reviewers' concerns. Our results show that current LLMs can generate detailed and fluent reviews but exhibit systematic weaknesses, such as overly positive recommendations, generic criticism, and uneven evidence grounding. We demonstrate that aggregate quality scores alone can overestimate review quality and argue for multi-dimensional evaluation of AI-generated peer reviews.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。