mR3实现72种语言的通用评分推理,比大模型更小更强。
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
- 构建跨语言评分模型,不依赖特定评分标准
- 在12种语言上超越更大模型,低资源语言也有效
- 适合多语言内容评估与偏好优化研究者使用
基于大语言模型的自动评估已在英语场景广泛采用并证明有效,但在非英语环境下表现不佳,且缺乏高效的多语言训练方法。本文提出mR3,一个覆盖72种语言的海量多语言、评分标准无关的奖励建模系统,为当前奖励建模领域最广的语言覆盖。通过全面研究数据与课程选择策略,识别出构建高质量奖励模型的有效方法,包括目标语言中的推理支持。mR3在多语言奖励模型基准测试中达到顶尖性能,优于更大模型(如GPT-OSS-120B)且体积最多缩小9倍。大量消融实验验证其有效性。进一步在离策略偏好优化中应用,并通过20名标注员在12种语言上的真人评测证实其推理过程与基于评分标准的评估质量高,即使对训练中未见的极低资源语言亦表现出色。相关模型、数据与代码已开源。
原文摘要 · Abstract (English)
Evaluation using Large Language Model (LLM) judges has been widely adopted in English and shown to be effective for automatic evaluation. However, their performance does not generalize well to non-English settings, and it remains unclear what constitutes effective multilingual training for such judges. In this paper, we introduce mR3, a massively multilingual, rubric-agnostic reward reasoning model trained on 72 languages, achieving the broadest language coverage in reward modeling to date. We present a comprehensive study of data and curriculum selection for training to identify effective strategies and data sources for building high-quality reward models, including support for reasoning in the target language. Our approach attains state-of-the-art performance on multilingual reward model benchmarks, surpassing much larger models (i.e., GPT-OSS-120B) while being up to 9x smaller, and its effectiveness is further confirmed through extensive ablation studies. Finally, we demonstrate the effectiveness of mR3 in off-policy preference optimization and validate the quality of its reasoning traces and rubric-based evaluations through human studies with 20 annotators across 12 languages, where mR3 models' reasoning is preferred, including for extremely low-resource languages that are entirely unseen during training. Our models, data, and code are available as open source at https://github.com/rubricreward/mr3.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。