让小模型模拟多人评审,高效精准评估大模型输出。
JudgePanel: A Compact Judge with Panel Deliberation via Adaptive Multi-Reward Reinforcement Learning

- 用多智能体讨论数据训练小模型,实现类似多人审议的判断能力。
- 140亿参数模型在4个评测集上超越700亿专用模型,且仅需数百样本快速适配新领域。
- 提出自适应奖励算法,动态调整不同评价目标的权重,提升评估稳定性。
LLM作为评判者范式已成为可扩展的人类评估替代方案。然而,单模型评判者受限于固有偏见,而通过多样化讨论缓解偏见的多智能体评估协议在推理时成本过高。为此,我们提出 extbf{ extit{JudgePanel}},为紧凑的 extit{Judge}模型赋予多智能体 extit{Panel}审议能力。首先,利用一组强评估器生成的审议轨迹进行训练,捕捉讨论、分歧与解决的结构化模式。为进一步提升判断质量,引入 extit{AdaReward}——一种自适应多奖励强化学习算法,能根据各目标饱和速率动态重平衡奖励权重。为支持实际部署,还设计了轻量级领域专化模块,仅需数百标注样本即可快速适应新评估领域。结果表明:(i) 创新性:首个在单模型推理成本下实现单模型多智能体审议能力的框架;(ii) 有效且可靠:140亿参数的JudgePanel在四个评估基准上超越高达700亿参数的专用模型,表现出强位置一致性,并可在数百样本下迅速适配新领域。
原文摘要 · Abstract (English)
The LLM-as-a-Judge paradigm has emerged as a scalable alternative to human evaluation. However, single-model judges are limited by their inherent model biases, while multi-agent evaluation protocols that mitigate this through diverse deliberation are prohibitively expensive at inference time. To this end, we propose \textbf{\modelname}, which equips a compact \underline{Judge} model with multi-agent \underline{Panel} deliberation capability. Specifically, we first train on panel deliberation traces from an ensemble of strong evaluators, capturing structured patterns of discussion, disagreement, and resolution. To further improve judgment quality beyond SFT, we introduce \textit{AdaReward}, an adaptive multi-reward RL algorithm that dynamically rebalances reward component weights as different objectives saturate at different rates during RL training. For practical deployment, we further design a lightweight domain specialization module for rapid adaptation to new evaluation domains with few hundred labeled samples. As a result, (i) \textit{Novel}: the first framework to equip a single compact judge with multi-agent panel deliberation capability at single-model inference cost; (ii) \textit{Effective \& Reliable}: JudgePanel with a 14B backbone outperforms judge-specialized models up to 70B across four evaluation benchmarks, demonstrates strong position consistency, and rapidly specializes to new domains with few hundred samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。