用少量示例提升低资源印地语奖励模型性能
Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples
- 通过检索高资源语言示例,增强低资源印地语的奖励信号
- 在博多语上准确率比零样本提示高12.81%
- 适合做多语言对齐但数据稀缺场景的研究者
奖励模型对大语言模型与人类偏好对齐至关重要。然而,大多数开源多语言奖励模型主要在高资源语言的偏好数据集上训练,导致低资源印地语的奖励信号不可靠。为低资源印地语收集大规模高质量偏好数据成本过高,基于偏好的训练方法不现实。为此,我们提出RELIC,一种针对低资源印地语的上下文学习奖励建模框架。RELIC通过成对排序目标训练一个检索器,从辅助高资源语言中选择最能区分优选与次优回答的上下文示例。在三个偏好数据集(PKU-SafeRLHF、WebGPT、HH-RLHF)上,使用先进开源奖励模型进行的大量实验表明,RELIC显著提升了低资源印地语奖励模型的准确性,始终优于现有示例选择方法。例如,在博多语(低资源印地语)上,使用LLaMA-3.2-3B奖励模型时,RELIC相比零样本提示准确率提升12.81%,相比最优示例选择方法提升10.13%。
原文摘要 · Abstract (English)
Reward models are essential for aligning large language models (LLMs) with human preferences. However, most open-source multilingual reward models are primarily trained on preference datasets in high-resource languages, resulting in unreliable reward signals for low-resource Indic languages. Collecting large-scale, high-quality preference data for these languages is prohibitively expensive, making preference-based training approaches impractical. To address this challenge, we propose RELIC, a novel in-context learning framework for reward modeling in low-resource Indic languages. RELIC trains a retriever with a pairwise ranking objective to select in-context examples from auxiliary high-resource languages that most effectively highlight the distinction between preferred and less-preferred responses. Extensive experiments on three preference datasets- PKU-SafeRLHF, WebGPT, and HH-RLHF-using state-of-the-art open-source reward models demonstrate that RELIC significantly improves reward model accuracy for low-resource Indic languages, consistently outperforming existing example selection methods. For example, on Bodo-a low-resource Indic language-using a LLaMA-3.2-3B reward model, RELIC achieves a 12.81% and 10.13% improvement in accuracy over zero-shot prompting and state-of-the-art example selection method, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。