arXiv:2503.23777cs.CL2025-03Conference of the …被引 1

通过过滤冲突梯度,提升多语言偏好对齐效果

CONGRAD:Conflicting Gradient Filtering for Multilingual Preference Alignment

  • 用梯度手术筛选跨语言冲突小的优质样本
  • 在10种语言上实现优于基线的对齐性能
  • 适合需要高效多语言模型对齐的研究者

大规模语言模型在多语言偏好对齐中进行联合训练时,常因目标冲突导致性能下降。现有研究虽关注多语言训练中的干扰问题,但对偏好对齐场景下的影响仍缺乏探索。为此,我们提出CONGRAD,一种可扩展且高效的样本过滤方法,通过梯度手术保留与聚合多语言更新方向一致的样本,减少跨语言梯度冲突。同时引入次线性梯度压缩策略,降低梯度累积过程中的内存开销。将CONGRAD集成至自奖励框架,在LLaMA3-8B和Gemma2-2B模型上,覆盖10种语言进行评估。结果表明,CONGRAD在已见与未见语言上均显著优于强基线,且对齐代价极低。

原文摘要 · Abstract (English)

Naive joint training of large language models (LLMs) for multilingual preference alignment can suffer from negative interference. This is a known issue in multilingual training, where conflicting objectives degrade overall performance. However, the impact of this phenomenon in the context of multilingual preference alignment remains largely underexplored. To address this issue, we propose CONGRAD, a scalable and effective filtering method that selects high-quality preference samples with minimal gradient conflicts across languages. Our method leverages gradient surgery to retain samples aligned with an aggregated multilingual update direction. Additionally, we incorporate a sublinear gradient compression strategy that reduces memory overhead during gradient accumulation. We integrate CONGRAD into self-rewarding framework and evaluate on LLaMA3-8B and Gemma2-2B across 10 languages. Results show that CONGRAD consistently outperforms strong baselines in both seen and unseen languages, with minimal alignment tax.

多语言偏好对齐梯度过滤大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。