打造通用判官模型,让AI能可靠评判各类任务。
CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards
- 用可验证奖励引导判断,通过拒绝采样提升推理能力。
- 7B小模型在多项评测中媲美2350亿参数大模型。
- 新增跨领域评测基准,统一评估标准适合研究者使用。
近期,以大语言模型为判官的评估方法受到关注。但现有判官模型存在领域狭窄、鲁棒性差的问题,难以全面评估。本文提出CompassJudger-2,一种通过任务驱动、多领域数据构建的通用判官模型。核心是采用可验证奖励监督判断任务,利用拒绝采样引导内在批判性推理,增强模型的鲁棒性和泛化能力。引入带有间隔策略的策略梯度损失函数优化学习目标。实验表明,CompassJudger-2在多个判官与奖励基准上表现优异,其7B模型在判断准确率上媲美DeepSeek-V3和Qwen3-235B-A22B等更大规模模型。此外,我们构建了JudgerBenchV2,一个涵盖跨领域判断准确率与排序一致性的综合性评测集,旨在标准化判官模型评估。该工作推动了可信赖、可扩展的大语言模型判别能力发展,并建立了新的性能与评估基准。
原文摘要 · Abstract (English)
Recently, the role of LLM-as-judge in evaluating large language models has gained prominence. However, current judge models suffer from narrow specialization and limited robustness, undermining their capacity for comprehensive evaluations. In this work, we present CompassJudger-2, a novel generalist judge model that overcomes these limitations via a task-driven, multi-domain data curation strategy. Central to our approach is supervising judgment tasks with verifiable rewards, guiding intrinsic critical reasoning through rejection sampling to foster robust, generalizable judgment capabilities. We introduce a refined learning objective with margin policy gradient loss to enhance performance. Empirically, CompassJudger-2 achieves superior results across multiple judge and reward benchmarks, and our 7B model demonstrates competitive judgment accuracy with significantly larger models like DeepSeek-V3 and Qwen3-235B-A22B. Additionally, we propose JudgerBenchV2, a comprehensive benchmark evaluating cross-domain judgment accuracy and rank consistency to standardize judge model evaluation. These contributions advance robust, scalable LLM judgment and establish new performance and evaluation standards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。