用英语安全能力提升多语言模型的安全性,避免性能下降。
MPO: Multilingual Safety Alignment via Reward Gap Optimization
- 利用英语模型的安全优势,优化其他语言的安全奖励差距
- 在三种大模型上验证,多语言安全性提升且通用能力不降
- 适合需要跨语言安全对齐的AI系统部署场景
大型语言模型(LLMs)在全球应用中日益重要,亟需可靠的多语言安全对齐以保障跨语言环境下的安全部署。现有基于偏好学习的安全对齐方法(如RLHF、DPO)主要针对单一语言,难以应对多语言数据中的噪声问题。为此,我们提出多语言奖励差距优化(MPO),利用主流语言(英语)已对齐的安全能力,优化目标语言与英语之间的奖励差距,从而有效迁移安全能力并保留主流语言原有优势。在LLaMA-3.1、Gemma-2和Qwen2.5三款大模型上的大量实验表明,MPO能有效实现多语言安全对齐,且未降低模型的多语言通用能力。
原文摘要 · Abstract (English)
Large language models (LLMs) have become increasingly central to AI applications worldwide, necessitating robust multilingual safety alignment to ensure secure deployment across diverse linguistic contexts. Existing preference learning methods for safety alignment, such as RLHF and DPO, are primarily monolingual and struggle with noisy multilingual data. To address these limitations, we introduce Multilingual reward gaP Optimization (MPO), a novel approach that leverages the well-aligned safety capabilities of the dominant language (English) to improve safety alignment across multiple languages. MPO directly minimizes the reward gap difference between the dominant language and target languages, effectively transferring safety capabilities while preserving the original strengths of the dominant language. Extensive experiments on three LLMs, LLaMA-3.1, Gemma-2 and Qwen2.5, validate MPO's efficacy in multilingual safety alignment without degrading general multilingual utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。