跨语言对比偏好调优让多语种模型无需标注也能提升表现
CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations
- 用自生成回复的对比学习优化多语言模型偏好
- 在14种语言上,多数任务表现超越基础模型
- 适合希望减少人工标注的多语种模型开发者
已有研究证明,通过奖励分数控制大语言模型自生成回复间的对比性,可提升英语下游偏好调优效果。本文将该方法扩展至多语言场景,在14种高资源与低资源语言上评估两个模型,涵盖多样化任务。核心发现:基于自生成的跨语言对比偏好调优(CroCo)可在无语言特定偏好标注的情况下实现迁移。在多语言基座上训练的英文奖励模型,能有效生成多数语言的内部排序,且单语或跨语言设置均优于各模型基础版本,同时避免监督微调导致的灾难性遗忘。结果表明,收益依赖在线策略数据;离线策略数据收益更高,而在线偏好优化未能超越离线方案。具体而言,在结构化任务上,EuroLLM-9B在6/7语言中达到或超过基线,Aya-3B在4/7设置中表现更优;在开放生成任务中,两种调优模型在11个语言上均胜过其对应基线。整体展示出多语言偏好调优的可行方向。
原文摘要 · Abstract (English)
Prior work establishes that controlled contrastiveness between self-generated responses from large language models, set via reward scores, improves downstream preference tuning in English. We extend this method to multiple languages and evaluate two models across a total of 14 high and low-resource languages on a diverse set of tasks. Our central finding is that cross-lingual contrastive preference tuning on self-generations (CroCo) transfers without language-specific preference annotation. A reward model trained on English preferences (atop a multilingual base) produces useful within-language rankings across most languages, and pairing in either a monolingual or multilingual setting improves over each model on the majority of setups while preventing the catastrophic forgetting of supervised fine-tuning. We observe that the gains require on-policy data. Off-policy responses reduce the benefit and online preference optimization fails to improve over the offline variant. Specifically, on structured tasks, our method matches or exceeds the base in 6/7 languages for EuroLLM-9B and 4/7 settings for Aya-3B. On open-ended generation, both tuned models win against their respective base across 11 evaluated languages. Overall, we show promising directions for multilingual preference tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。