arXiv:2509.04484cs.CLcs.AI2025-09EMNLP综述被引 19

自动评估审稿意见对作者的帮助程度,提升评审效率与质量

The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors

  • 定义审稿反馈四维度:可操作性、具体性、可验证性、帮助性
  • 构建1430条人工标注+1万条合成标注数据集,含评分理由
  • 模型表现接近甚至超越GPT-4o,但仍不及人类审稿质量

为作者提供有建设性的反馈是同行评审的核心。随着审稿人时间日益紧张,亟需自动化支持系统保障评审质量,使反馈真正对作者有用。为此,我们识别出四个决定审稿意见效用的关键方面:可操作性、具体性与可验证性、帮助性。为促进相关模型的评估与发展,我们引入了RevUtil数据集,包含1,430条人工标注的审稿评论,并通过10,000条合成标注数据用于训练。合成数据还包含理由(rationales),即每条评论得分的解释。基于该数据集,我们对微调模型在评估评论各维度及生成理由方面进行基准测试。实验表明,这些微调模型与人类的一致性水平相当,甚至在某些情况下超过强大的闭源模型如GPT-4o。进一步分析显示,机器生成的评审普遍在四项指标上弱于人类评审。

原文摘要 · Abstract (English)

Providing constructive feedback to paper authors is a core component of peer review. With reviewers increasingly having less time to perform reviews, automated support systems are required to ensure high reviewing quality, thus making the feedback in reviews useful for authors. To this end, we identify four key aspects of review comments (individual points in weakness sections of reviews) that drive the utility for authors: Actionability, Grounding & Specificity, Verifiability, and Helpfulness. To enable evaluation and development of models assessing review comments, we introduce the RevUtil dataset. We collect 1,430 human-labeled review comments and scale our data with 10k synthetically labeled comments for training purposes. The synthetic data additionally contains rationales, i.e., explanations for the aspect score of a review comment. Employing the RevUtil dataset, we benchmark fine-tuned models for assessing review comments on these aspects and generating rationales. Our experiments demonstrate that these fine-tuned models achieve agreement levels with humans comparable to, and in some cases exceeding, those of powerful closed models like GPT-4o. Our analysis further reveals that machine-generated reviews generally underperform human reviews on our four aspects.

同行评审自动评估数据集自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。