多语言拒答对齐提升大模型跨语安全,避免有害内容漏判。
Multilingual Refusal Alignment for Safer Large Language Models

- 构建12种欧洲语言的拒答对齐数据集RefusEU,支持跨语言评估。
- 仅英语对齐无法保证多语言安全,同一危害类别在不同语言中表现不一。
- 多语言训练可同时提升安全性与通用能力,不影响模型整体表现。
随着大语言模型在全球部署,确保其在多种语言中的安全性和对齐性至关重要。然而,安全行为在不同语言间常表现出不可预测的差异,带来一致且伦理化的AI挑战。本文系统研究多语言对齐机制,探讨单语言对齐是否能跨语言迁移、训练过程中语言一致性如何保持,以及与通用知识能力之间的权衡。我们提出RefusEU,一个覆盖12种欧洲语言的拒答对齐数据集,包含专门用于评估当前顶尖模型的测试集。通过受控的直接偏好优化(DPO)实验,发现仅在英语中进行对齐不足以确保跨语言安全,即使针对相同危害类别;而使用多语言数据集训练则能在不损害通用性能的前提下提升安全性,该结论基于全球MMLU基准测试验证。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) are deployed globally, ensuring their safety and alignment across multiple languages becomes paramount. However, safety behaviors often vary unpredictably between languages, posing significant challenges for consistent and ethical AI. In this work, we systematically investigate the dynamics of multilingual alignment, exploring whether single-language alignment transfers cross-lingually, how language consistency is preserved during training, and the resulting trade-offs with general knowledge capabilities. We introduce RefusEU, a novel refusal alignment dataset covering 12 European languages, including a dedicated test set for evaluating current state-of-the-art models. Our controlled Direct Preference Optimization (DPO) experiments provide two key insights: aligning models exclusively in English is insufficient to ensure cross-lingual safety, even for the same harm categories, whereas training on multilingual datasets can improve safety without degrading general performance, as measured by the Global MMLU benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。