构建首个跨学科多模态审稿数据集,支持图文结合的智能审稿研究。
FMMD: A multimodal open peer review dataset based on F1000Research
- 从F1000Research收集多版本论文与审稿意见,实现评论与具体版本精确对齐。
- 包含图文结构数据,突破传统文本主导的审稿数据局限。
- 适合研究智能审稿、多模态分析及跨领域同行评议的学者使用。
自动化学术论文评审(ASPR)已进入与传统同行评审并行阶段,人工智能系统正越来越多地融入真实稿件评估流程。与此同时,自动化与AI辅助审稿研究迅猛发展。然而,现有数据集存在诸多关键局限:审稿人常需评估图表、表格及复杂版式以判断科学主张,但多数数据集仍高度偏向文本;且数据主要来自计算机科学会议,覆盖面窄;此外,缺乏审稿意见与具体稿件版本间的精确对应,难以揭示审稿过程与稿件演进的动态关系。为此,我们推出FMMD,一个基于F1000Research构建的多模态、跨学科开放审稿数据集。该数据集整合了稿件级别的视觉与结构数据,以及版本特定的审稿报告与编辑决策。通过明确对齐审稿意见与具体文章迭代版本,FMMD支持对跨学科科研领域中同行评审生命周期的细粒度分析。该数据集可支撑多模态问题检测与多模态审稿意见生成等任务,为同行评审研究提供全面的实证资源。
原文摘要 · Abstract (English)
Automated scholarly paper review (ASPR) has entered the coexistence phase with traditional peer review, where artificial intelligence (AI) systems are increasingly incorporated into real-world manuscript evaluation. In parallel, research on automated and AI-assisted peer review has proliferated. Despite this momentum, empirical progress remains constrained by several critical limitations in existing datasets. While reviewers routinely evaluate figures, tables, and complex layouts to assess scientific claims, most existing datasets remain overwhelmingly text-centric. This bias is reinforced by a narrow focus on data from computer science venues. Furthermore, these datasets lack precise alignment between reviewer comments and specific manuscript versions, obscuring the iterative relationship between peer review and manuscript evolution. In response, we introduce FMMD, a multimodal and multidisciplinary open peer review dataset curated from F1000Research. The dataset bridges the current gap by integrating manuscript-level visual and structural data with version-specific reviewer reports and editorial decisions. By providing explicit alignment between reviewer comments and the exact article iteration under review, FMMD enables fine-grained analysis of the peer review lifecycle across diverse scientific domains. FMMD supports tasks such as multimodal issue detection and multimodal review comment generation. It provides a comprehensive empirical resource for the development of peer review research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。