arXiv:2609.05910cs.CLcs.AI2026-09被引 1

统一多语言与评估范式的推理型奖励模型,提升开放任务评价可靠性。

UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation Paradigms

论文配图:UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation Paradigms
图 1 · 摘自论文原文
  • 采用分阶段推理链动态生成评价标准,适应不同指令与输入。
  • 在多语言、多任务基准上性能接近同规模顶尖模型。
  • 支持未见评估范式,适合需要跨语言公平评估的场景。

强化学习在可验证奖励任务中表现优异,但在开放性任务中,奖励模型的可靠性仍是关键挑战。现有方案或依赖昂贵的专用大模型评分系统,或使用不可解释的标量奖励模型。近期生成式奖励模型虽具潜力,但仍受限于静态评价标准、碎片化评估范式及有限的多语言支持。为此,我们构建了覆盖六大领域、103种语言的大规模多语言数据集 MixReward,包含成对与列表数据,并提出 UniRRM——一种支持多语言与多种评估范式的统一推理奖励模型。UniRRM 通过分阶段推理链动态生成通用任务与指令特定的评价标准,实现细粒度、输入自适应判断,同时保持跨语言一致性。实验表明,UniRRM-8B 与 UniRRM-14B 在多个基准测试中表现接近同规模最优模型,且在未见过的评估范式下仍有效。消融实验验证了其可靠性和有效性。

原文摘要 · Abstract (English)

Reinforcement learning (RL) excels on tasks with verifiable rewards, but in open-ended tasks, the reliability of reward models remains a key challenge. Existing solutions either depend on costly proprietary LLM-as-a-Judge systems or opaque scalar reward models that lack interpretability. Recent works on generative reward models offer a promising alternative, but they remain constrained by static evaluation criteria, fragmented evaluation paradigms, and limited multilingual support. To address these challenges, we introduce \textbf{MixReward}, a large-scale multilingual dataset spanning six domains and 103 languages, containing both pairwise and listwise data, and propose \textbf{UniRRM}, a unified reasoning reward model supporting multiple languages and evaluation paradigms. UniRRM uses a staged reasoning chain to dynamically generate task-generic and instruction-specific criteria, enabling fine-grained, input-adaptive judgments while maintaining consistency across languages. Experiments demonstrate that UniRRM-8B and UniRRM-14B achieve performance close to the state-of-the-art for models of comparable size across multiple benchmarks, and are effective for unseen evaluation paradigms. In addition, ablation studies validate the reliability and effectiveness of UniRRM.

奖励模型多语言强化学习推理链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。