提出评估AI生成审稿意见的框架,提升审稿质量与可信度
ReviewEval: An Evaluation Framework for AI-Generated Reviews
- 构建ReviewEval框架,多维度评估AI审稿的准确性与专业性
- 自研ReviewAgent使建议可操作性提升47.62%,分析深度增12.73%
- 适合需要高效高质量审稿的学术机构与期刊编辑
随着学术研究数量激增与合格审稿人短缺,亟需创新方法应对同行评审挑战。本文提出:1. ReviewEval——一套全面评估AI生成审稿意见的框架,涵盖与人类判断的一致性、事实准确性、分析深度、建设性程度及对审稿指南的遵循情况;2. ReviewAgent——基于大模型的审稿生成代理,具备针对目标会议/期刊的对齐机制、自我优化循环以及通过ReviewEval进行外部改进的增强环路。实验表明,ReviewAgent在可操作性上比现有AI基线提升6.78%,比专家审稿提升47.62%;分析深度提升3.97%和12.73%;对审稿指南的遵守率分别提高10.11%和47.26%。该工作确立了AI辅助同行评审的关键指标,显著提升了AI生成审稿意见的可靠性与影响力。
原文摘要 · Abstract (English)
The escalating volume of academic research, coupled with a shortage of qualified reviewers, necessitates innovative approaches to peer review. In this work, we propose: 1. ReviewEval, a comprehensive evaluation framework for AI-generated reviews that measures alignment with human assessments, verifies factual accuracy, assesses analytical depth, identifies degree of constructiveness and adherence to reviewer guidelines; and 2. ReviewAgent, an LLM-based review generation agent featuring a novel alignment mechanism to tailor feedback to target conferences and journals, along with a self-refinement loop that iteratively optimizes its intermediate outputs and an external improvement loop using ReviewEval to improve upon the final reviews. ReviewAgent improves actionable insights by 6.78% and 47.62% over existing AI baselines and expert reviews respectively. Further, it boosts analytical depth by 3.97% and 12.73%, enhances adherence to guidelines by 10.11% and 47.26% respectively. This paper establishes essential metrics for AIbased peer review and substantially enhances the reliability and impact of AI-generated reviews in academic research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。