首个多语言奖励模型评估基准,揭示英外语言性能差距。
M-RewardBench: Evaluating Reward Models in Multilingual Settings
- 构建23种语言的2870个偏好样本,覆盖对话、安全、推理等能力
- 发现非英语语言上奖励模型性能显著下降,跨语言偏好差异大
- 翻译质量与资源丰富度影响模型表现,适合多语言研究者参考
奖励模型(RMs)通过融入人类反馈推动了当前大语言模型的性能提升。然而,现有研究主要聚焦于英文场景,对多语言环境下RMs的表现仍缺乏系统评估。本文首次构建多语言奖励模型评估基准M-RewardBench,包含2.87千个偏好实例,覆盖23种类型多样的语言,用于测试模型在对话、安全、推理和翻译方面的表现。我们对多种奖励模型在该基准上的表现进行了严谨评估,发现其在非英语语言中的性能显著低于英语,且不同语言间偏好存在显著差异。同时,我们验证了翻译质量与资源丰富度对模型性能的正向影响。研究结果表明,高质量翻译能提升奖励模型表现,高资源语言更易获得准确偏好判断。本文发布M-RewardBench数据集及代码库,以促进多语言奖励模型评估的发展。
原文摘要 · Abstract (English)
Reward models (RMs) have driven the state-of-the-art performance of LLMs today by enabling the integration of human feedback into the language modeling process. However, RMs are primarily trained and evaluated in English, and their capabilities in multilingual settings remain largely understudied. In this work, we conduct a systematic evaluation of several reward models in multilingual settings. We first construct the first-of-its-kind multilingual RM evaluation benchmark, M-RewardBench, consisting of 2.87k preference instances for 23 typologically diverse languages, that tests the chat, safety, reasoning, and translation capabilities of RMs. We then rigorously evaluate a wide range of reward models on M-RewardBench, offering fresh insights into their performance across diverse languages. We identify a significant gap in RMs' performances between English and non-English languages and show that RM preferences can change substantially from one language to another. We also present several findings on how different multilingual aspects impact RM performance. Specifically, we show that the performance of RMs is improved with improved translation quality. Similarly, we demonstrate that the models exhibit better performance for high-resource languages. We release M-RewardBench dataset and the codebase in this study to facilitate a better understanding of RM evaluation in multilingual settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。