arXiv:2602.08896cs.IRcs.AI2026-02被引 1

构建大规模真实评审推荐数据集并提出新型模型,提升专家匹配精度。

OmniReview: A Large-scale Benchmark and LLM-enhanced Framework for Realistic Reviewer Recommendation

  • 用多源平台数据构建20万+条验证过的评审记录,支持真实场景评估。
  • 提出融合LLM与多任务学习的新框架,在7项指标中6项达顶尖水平。
  • 适合需要高精度、可解释性评审推荐的学术出版与AI研究者使用。

学术同行评审仍是学术验证的核心,但当前研究受限于数据规模小、评估指标简化,难以反映真实编辑流程。为此,我们构建了OmniReview,一个整合多源学术平台的综合性数据集,通过消歧流程获得202,756条经验证的评审记录。基于此,提出三层级评估框架,从召回率到精准专家识别进行系统评估。在方法层面,现有基于嵌入的方法存在语义压缩瓶颈且可解释性差。为此,我们提出基于大语言模型的多门控专家混合模型(Pro-MMoE),利用LLM生成的语义档案保留细粒度专长特征,并结合任务自适应的MMoE架构动态平衡冲突目标。大量实验表明,Pro-MMoE在七项指标中有六项达到领先水平,确立了真实评审推荐的新基准。

原文摘要 · Abstract (English)

Academic peer review remains the cornerstone of scholarly validation, yet the field faces some challenges in data and methods. From the data perspective, existing research is hindered by the scarcity of large-scale, verified benchmarks and oversimplified evaluation metrics that fail to reflect real-world editorial workflows. To bridge this gap, we present OmniReview, a comprehensive dataset constructed by integrating multi-source academic platforms encompassing comprehensive scholarly profiles through the disambiguation pipeline, yielding 202, 756 verified review records. Based on this data, we introduce a three-tier hierarchical evaluaion framework to assess recommendations from recall to precise expert identification. From the method perspective, existing embedding-based approaches suffer from the information bottleneck of semantic compression and limited interpretability. To resolve these method limitations, we propose Profiling Scholars with Multi-gate Mixture-of-Experts (Pro-MMoE), a novel framework that synergizes Large Language Models (LLMs) with Multi-task Learning. Specifically, it utilizes LLM-generated semantic profiles to preserve fine-grained expertise nuances and interpretability, while employing a Task-Adaptive MMoE architecture to dynamically balance conflicting evaluation goals. Comprehensive experiments demonstrate that Pro-MMoE achieves state-of-the-art performance across six of seven metrics, establishing a new benchmark for realistic reviewer recommendation.

评审推荐LLM应用多任务学习学术数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。