英语奖励模型可跨语言迁移,显著提升多语言对齐效果
Cross-lingual Transfer of Reward Models in Multilingual Alignment
- 用英语训练的奖励模型直接迁移到其他语言
- 在多语言评估中表现优于本地训练模型3~4%
- 适合做多语言大模型对齐的研究者和开发者
基于人类反馈的强化学习(RLHF)高度依赖精确的奖励模型(RMs)。然而,现有奖励建模研究主要集中于英语,限制了RLHF在多语言对齐中的应用。本文研究了从英语等语言跨语言迁移奖励模型的能力。实验表明,英语训练的奖励模型在目标语言上表现优异,在Multilingual RewardBench上平均得分比本地训练模型高出3~4%。我们进一步分析了跨语言迁移过程中的表征偏移问题,并通过多语言对齐任务展示了奖励模型迁移如何提升指令跟随能力。本文还对现成奖励模型进行了广泛分析,并开源代码、模型与数据。
原文摘要 · Abstract (English)
Reinforcement learning with human feedback (RLHF) is shown to largely benefit from precise reward models (RMs). However, recent studies in reward modeling schemes are skewed towards English, limiting the applicability of RLHF in multilingual alignments. In this work, we investigate the cross-lingual transfer of RMs trained in diverse languages, primarily from English. Our experimental results demonstrate the strong cross-lingual transfer of English RMs, exceeding target language RMs by 3~4% average increase in Multilingual RewardBench. Furthermore, we analyze the cross-lingual transfer of RMs through the representation shifts. Finally, we perform multilingual alignment to exemplify how cross-lingual transfer in RM propagates to enhanced multilingual instruction-following capability, along with extensive analyses on off-the-shelf RMs. We release the code, model, and data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。