发现大模型对低资源语言的安全对齐存在明显差距,需针对性优化。
The Hidden Space of Safety: Understanding Preference-Tuned LLMs in Multilingual context
- 通过分析对齐前后嵌入空间变化,揭示多语言安全对齐机制差异。
- 7个模型在高低资源语言间安全表示差异显著,低资源语言对齐效果差。
- 适合关注多语言AI公平性与安全性的研究者与开发者参考。
对齐训练使大语言模型在推理、指令遵循和减少有害生成方面表现优异,但其广泛部署中仍存在明显的单语偏差,引发对跨语言对齐有效性的问题。当前对齐方法主要针对英语,尚不清楚其在多语言场景下的泛化能力。为此,我们系统分析了对齐前后大模型嵌入空间的分布偏移,评估其对不同语言行为的影响。利用对齐导致的安全空间分离作为量化工具,衡量对齐对安全约束的强制程度。研究采用平衡毒性数据集和并行文本净化基准,评估了7个大语言模型,发现高资源语言与低资源语言在潜在表征空间中存在显著差异。这些结果凸显了为不同语言进行特定微调的必要性,以确保多语言对齐的公平、可靠与稳健。本研究为构建真正安全的多语言大模型提供了基础,强调了亟需解决低资源语言中的对齐鸿沟问题。
原文摘要 · Abstract (English)
Alignment tuning has enabled large language models to excel in reasoning, instruction-following, and minimizing harmful generations. However, despite their widespread deployment, these models exhibit a monolingual bias, raising concerns about the effectiveness of alignment across languages. Current alignment methods predominantly focus on English, leaving it unclear how alignment mechanism generalize to multilingual settings. To address this, we conduct a systematic analysis of distributional shifts in the embedding space of LLMs before and after alignment, uncovering its impact on model behavior across diverse languages. We leverage the alignment-induced separation in safety space as a quantitative tool to measure how alignment enforces safety constraints. Our study evaluates seven LLMs using balanced toxicity datasets and parallel text-detoxification benchmarks, revealing substantial disparities in the latent representation space between high-resource and low-resource languages. These findings underscore the need for language-specific fine-tuning to ensure fair, reliable and robust multilingual alignment. Our insights provide a foundation for developing truly safe multilingual LLMs, emphasizing the urgency of addressing alignment gaps in underrepresented languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。