arXiv:2601.19637cs.CL2026-01ACL综述

构建新基准与框架,提升论文评审专家匹配精度。

RATE: Reviewer Profiling and Annotation-free Training for Expertise Ranking in Peer Review Systems

  • 基于2024-2025年论文构建高保真评审画像数据集
  • 通过关键词摘要与弱监督训练实现精准专家排序
  • 适合需要高效、准确评审分配的会议或期刊

在大模型时代,论文评审分配面临严峻挑战:快速演进的研究主题使许多2023年前的评估基准过时,且传统代理信号无法真实反映评审人熟悉度。为此,我们构建了LR-bench——一个基于2024-2025年人工智能/自然语言处理论文的高保真、最新基准,通过大规模邮件调查获取五级自评熟悉度,共获得1055个专家标注的论文-评审人评分样本。我们进一步提出RATE框架,将每位评审人近期论文提炼为紧凑的关键词画像,并利用启发式检索信号构建弱偏好监督,微调嵌入模型,实现论文与评审人画像的直接匹配。在LR-bench和CMU黄金标准数据集上,该方法持续达到顶尖性能,显著优于现有嵌入基线。相关数据集已开源至https://huggingface.co/datasets/Gnociew/LR-bench,代码库见https://github.com/Gnociew/RATE-Reviewer-Assign。

原文摘要 · Abstract (English)

Reviewer assignment is increasingly critical yet challenging in the LLM era, where rapid topic shifts render many pre-2023 benchmarks outdated and where proxy signals poorly reflect true reviewer familiarity. We address this evaluation bottleneck by introducing LR-bench, a high-fidelity, up-to-date benchmark curated from 2024-2025 AI/NLP manuscripts with five-level self-assessed familiarity ratings collected via a large-scale email survey, yielding 1055 expert-annotated paper-reviewer-score annotations. We further propose RATE, a reviewer-centric ranking framework that distills each reviewer's recent publications into compact keyword-based profiles and fine-tunes an embedding model with weak preference supervision constructed from heuristic retrieval signals, enabling matching each manuscript against a reviewer profile directly. Across LR-bench and the CMU gold-standard dataset, our approach consistently achieves state-of-the-art performance, outperforming strong embedding baselines by a clear margin. We release LR-bench at https://huggingface.co/datasets/Gnociew/LR-bench, and a GitHub repository at https://github.com/Gnociew/RATE-Reviewer-Assign.

评审分配专家匹配嵌入模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。