构建跨学科多模态论文评审基准,评估大模型自动审稿能力。
MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation
- 设计13项任务覆盖审稿全流程,涵盖图文混合内容
- 包含240篇跨四大领域的论文与专家评审意见
- 适合研究AI辅助学术评审的开发者和评测人员
随着学术论文数量激增,同行评审已成为研究社区不可或缺但耗时的任务。尽管大语言模型(LLMs)被用于生成评审意见,但当前缺乏统一的评估基准来严格检验模型在生成全面、准确且符合人类判断的评审方面的表现,尤其在涉及图表等多模态内容时。为此,我们提出 extbf{MMReview},一个涵盖多个学科和模态的综合性基准。该基准包含来自人工智能、自然科学、工程科学和社会科学四大领域共17个研究方向的240篇论文,每篇均配有专家撰写的评审意见及多模态内容。我们设计了13项任务,分为四个核心类别:逐步生成评审、结果归纳、与人类偏好对齐,以及对抗性输入下的鲁棒性测试。在16个开源模型和5个闭源先进模型上的实验验证了该基准的全面性。我们期望MMReview能成为构建标准化自动同行评审系统的关键一步。
原文摘要 · Abstract (English)
With the rapid growth of academic publications, peer review has become an essential yet time-consuming responsibility within the research community. Large Language Models (LLMs) have increasingly been adopted to assist in the generation of review comments; however, current LLM-based review tasks lack a unified evaluation benchmark to rigorously assess the models' ability to produce comprehensive, accurate, and human-aligned assessments, particularly in scenarios involving multimodal content such as figures and tables. To address this gap, we propose \textbf{MMReview}, a comprehensive benchmark that spans multiple disciplines and modalities. MMReview includes multimodal content and expert-written review comments for 240 papers across 17 research domains within four major academic disciplines: Artificial Intelligence, Natural Sciences, Engineering Sciences, and Social Sciences. We design a total of 13 tasks grouped into four core categories, aimed at evaluating the performance of LLMs and Multimodal LLMs (MLLMs) in step-wise review generation, outcome formulation, alignment with human preferences, and robustness to adversarial input manipulation. Extensive experiments conducted on 16 open-source models and 5 advanced closed-source models demonstrate the thoroughness of the benchmark. We envision MMReview as a critical step toward establishing a standardized foundation for the development of automated peer review systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。