构建大规模审稿人推荐数据集,助力学术出版高效化
FRONTIER-RevRec: A Large-scale Dataset for Reviewer Recommendation
- 基于2007-2025年真实审稿记录构建大规模数据集
- 内容匹配方法显著优于协同过滤,语言模型更擅长捕捉语义对齐
- 适合研究学术推荐、同行评审系统优化的学者使用
审稿人推荐是提升学术出版效率的关键任务,但长期受限于高质量基准数据集的缺乏。为此,我们提出FRONTIER-RevRec,一个基于Frontiers开放获取平台(2007–2025)真实同行评审记录构建的大规模数据集。该数据集涵盖209个期刊、478,379篇论文和177,941位独特审稿人,覆盖临床医学、生物、心理、工程与社会科学等多个领域。在该数据集上的综合评估表明,内容型方法显著优于协同过滤。结构分析揭示了学术推荐与商业推荐的根本差异。特别地,基于语言模型的方法在捕捉论文内容与审稿人专长间的语义对齐方面表现突出。此外,实验识别出优化推荐流程的最佳聚合策略。FRONTIER-RevRec旨在成为推动审稿人推荐研究的综合性基准,助力构建更高效的学术同行评审系统。数据集已公开:https://anonymous.4open.science/r/FRONTIER-RevRec-5D05。
原文摘要 · Abstract (English)
Reviewer recommendation is a critical task for enhancing the efficiency of academic publishing workflows. However, research in this area has been persistently hindered by the lack of high-quality benchmark datasets, which are often limited in scale, disciplinary scope, and comparative analyses of different methodologies. To address this gap, we introduce FRONTIER-RevRec, a large-scale dataset constructed from authentic peer review records (2007-2025) from the Frontiers open-access publishing platform https://www.frontiersin.org/. The dataset contains 177941 distinct reviewers and 478379 papers across 209 journals spanning multiple disciplines including clinical medicine, biology, psychology, engineering, and social sciences. Our comprehensive evaluation on this dataset reveals that content-based methods significantly outperform collaborative filtering. This finding is explained by our structural analysis, which uncovers fundamental differences between academic recommendation and commercial domains. Notably, approaches leveraging language models are particularly effective at capturing the semantic alignment between a paper's content and a reviewer's expertise. Furthermore, our experiments identify optimal aggregation strategies to enhance the recommendation pipeline. FRONTIER-RevRec is intended to serve as a comprehensive benchmark to advance research in reviewer recommendation and facilitate the development of more effective academic peer review systems. The FRONTIER-RevRec dataset is available at: https://anonymous.4open.science/r/FRONTIER-RevRec-5D05.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。