arXiv:2605.27865cs.CL2026-05

用评分标准指导训练,自动匹配论文与合适审稿人。

MERIT: Matching Expertise via Rubric-Informed Training for Reviewer Assignment

论文配图:MERIT: Matching Expertise via Rubric-Informed Training for Reviewer Assignment
图 1 · 摘自论文原文
  • 用强化学习训练评审评估器,根据论文需求匹配审稿人专长。
  • 40亿参数模型比更大通用模型更准,检索器性能领先基准数据集。
  • 适合大规模会议审稿分配,无需人工标注,可快速部署。

大规模会议面临论文与审稿人匹配的挑战,现有方法或依赖粗粒度代理信号,混淆相关性与真正适配性,或需昂贵的人工标注难以扩展。我们提出MERIT,一种两阶段框架,将基于标准的专长匹配转化为可扩展的适配性监督。第一阶段通过强化学习训练一个评审评估器,识别论文所需的专长维度,匹配审稿人过往工作,并生成适配性判断,奖励由基于论文特定评标标准的LLM裁判提供。第二阶段将评估器预测蒸馏为基于嵌入的检索器,实现高效大规模分配。实验表明,40亿参数的评审评估器在适配性分类上优于更大通用模型,所生成的检索器在LR-Bench和CMU Gold数据集上达到当前最佳表现。代码已开源。

原文摘要 · Abstract (English)

Matching submissions with suitable reviewers at scale is a growing challenge for major venues, yet existing approaches either rely on coarse proxy signals that conflate general relatedness with true suitability, or require expensive human annotations that are difficult to scale for training. We propose MERIT, a two-stage framework that bridges this gap by converting criterion-level expertise matching into scalable suitability supervision. In the first stage, we train a reviewer assessor via reinforcement learning to identify the expertise dimensions a paper requires, match them against the reviewer's prior work, and produce a suitability decision, with rewards provided by an LLM judge guided by paper-specific expertise rubrics. In the second stage, we distill the assessor's predictions into an embedding-based retriever for efficient large-scale assignment. Experiments show that our 4B reviewer assessor outperforms larger general-purpose LLMs on suitability classification, and the resulting retriever achieves state-of-the-art performance across LR-Bench and the CMU Gold dataset. Our code is available at https://github.com/Luli3220/MERIT.

审稿分配强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。