提升多语言模型跨语种响应一致性,让翻译等价提示在不同语言下表现更一致。
Post-Training Language Models for Crosslingual Consistency
- 提出直接一致性优化(DCO),通过离策略训练提升跨语言响应一致性。
- 在26种语言上实验,显著优于现有方法,尤其提升低资源语言对齐效果。
- 适用于需要跨语言行为一致性的场景,如多语言问答与翻译系统。
语言模型在不同语言间对翻译等价提示的回应常不一致,影响多语言系统的可靠性。本文从信息论角度定义跨语言一致性为模型输出分布与其跨语言往返映射之间的差异上界。提出惩罚式一致性优化(PCO),通过引入Kullback-Leibler散度项,将模型与固定参考语言模型对齐。由于直接优化需昂贵的在线采样,我们设计可离策略优化的近似方法——直接一致性优化(DCO)。在26种语言、多种语言模型上的实验表明,DCO显著提升跨语言一致性,优于现有方法,并能有效实现低资源语言的定向对齐。
原文摘要 · Abstract (English)
Language models often respond inconsistently to translation-equivalent prompts across languages, undermining the reliability of multilingual systems. To quantify this, we give an information-theoretic definition of crosslingual consistency as a divergence bound between a model's response distribution and its round-trip pushforward across languages. We then introduce penalized consistency optimization (PCO), a post-training procedure that couples this divergence with a Kullback-Leibler penalty to a fixed reference language model. Because direct optimization of PCO requires expensive on-policy roll-outs, we propose a tractable surrogate, direct consistency optimization (DCO), which can be optimized off-policy. Across diverse language models and 26 languages, DCO significantly improves crosslingual consistency, outperforms existing methods, and enables targeted alignment of low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。