arXiv:2602.16660cs.CLcs.AI2026-02中稿 · ICLR被引 8

一次对齐,多语受益:用统一方法提升多语言模型安全对齐

Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment

  • 引入可插拔的多语言一致性损失,单次更新实现跨语言语义一致
  • 仅用多语言提示变体,无需低资源语言的额外响应监督
  • 适配多种模型架构,兼顾安全对齐与通用能力,适合多语言部署

大型语言模型在多语言社区中的广泛应用,要求可靠的多语言安全对齐。然而,现有方法扩展到其他语言时往往需要大量资源,或依赖高资源语言的成对对齐,限制了可扩展性。本文提出一种资源高效的多语言安全对齐方法,引入即插即用的多语言一致性(MLC)损失,可集成至现有单语言对齐流程中。通过增强多语言表征向量间的共线性,该方法在单次更新中促进跨语言语义层面的方向一致性。由此可在不需低资源语言额外响应监督的前提下,同时实现多语言对齐。我们在不同模型架构和对齐范式下验证了该方法,结果表明其在有限影响通用性能的情况下有效提升多语言安全性。跨语言任务评估显示更强的泛化能力,证明该方法在有限监督下具有实用价值。

原文摘要 · Abstract (English)

The widespread deployment of large language models (LLMs) across linguistic communities necessitates reliable multilingual safety alignment. However, recent efforts to extend alignment to other languages often require substantial resources, either through large-scale, high-quality supervision in the target language or through pairwise alignment with high-resource languages, which limits scalability. In this work, we propose a resource-efficient method for improving multilingual safety alignment. We introduce a plug-and-play Multi-Lingual Consistency (MLC) loss that can be integrated into existing monolingual alignment pipelines. By improving collinearity between multilingual representation vectors, our method encourages directional consistency at the multilingual semantic level in a single update. This allows simultaneous alignment across multiple languages using only multilingual prompt variants without requiring additional response-level supervision in low-resource languages. We validate the proposed method across different model architectures and alignment paradigms, and demonstrate its effectiveness in enhancing multilingual safety with limited impact on general model utility. Further evaluation across languages and tasks indicates improved cross-lingual generalization, suggesting the proposed approach as a practical solution for multilingual consistency alignment under limited supervision.

多语言对齐安全对齐高效训练扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。