arXiv:2607.13315cs.CL2026-07

用少样本数据让多语言大模型更好理解人类偏好

Meta-Learning Preferences for Multilingual LLM Alignment

论文配图:Meta-Learning Preferences for Multilingual LLM Alignment
图 1 · 摘自论文原文
  • 通过跨语言元学习,从多种语言数据中提取可迁移的偏好初始化
  • 仅用100条目标语言数据,胜率提升最高达28%
  • 适合低资源语言对齐,支持不同语言组合与模型规模

不同语言间人类偏好数据分布不均,给多语言大模型对齐带来挑战。为解决低资源语言对齐的数据不足问题,我们提出一种面向强化学习与直接偏好优化的元学习框架。该框架利用其他语言的偏好数据,学习可迁移的初始化,使目标语言在极少数据下仍能有效适配。我们为元奖励建模与元策略优化提供了理论保障,并在多语言基准上实证验证了方法有效性。在仅100条目标语言偏好样本的极低资源设置下,本方法相比基线最高提升28%胜率,且在多个目标语言和模型规模上持续领先。优势在不同元训练语言组合及语言距离下均保持稳定。

原文摘要 · Abstract (English)

Unequal availability of human preference data across languages poses a significant challenge for aligning large language models in multilingual settings. To address the lack of sufficient data in low-resource language alignment, we propose a meta-learning framework for Reinforcement Learning from Human Feedback and Direct Preference Optimization. By leveraging preference data from other languages, our framework learns a transferable initialization that enables effective adaptation to a target language with minimal data. We provide theoretical guarantees for both the meta-reward modeling and meta-policy optimization settings, and empirically demonstrate the effectiveness of our approach on multilingual benchmarks. In an extremely low-resource setting with only 100 target-language preference samples, our approach achieves up to $28\%$ win-rate improvements over baseline methods, and consistently outperforms baselines across multiple target languages and model scales. Our approaches retain these advantages across different combinations of meta-training languages and varying linguistic distances from the target languages.

多语言偏好对齐元学习低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。