低资源语言的LLM安全对齐存在明显短板,这篇综述系统梳理了问题与解决方案。
LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review

- 按数据、目标、机制三类归纳安全对齐方法,构建分类体系
- 发现多语言模型易受跨语言攻击,非洲等地区缺乏安全基准
- 强调需文化适配的评估框架和参与式数据收集,适合安全研究者参考
大型语言模型(LLMs)在安全对齐方面取得显著进展,但在低资源和多语言环境下其安全保障远弱于高资源语言。本文采用PRISMA 2020方法,对来自Semantic Scholar、arXiv和OpenAlex的约1500篇论文进行系统文献综述,筛选出50篇相关研究进行分析。综述围绕四大主题展开:安全对齐方法、多语言安全风险、评估基准与跨语言可迁移性。提出基于数据适应、目标优化和机制对齐三类适应机制的安全对齐分类体系。研究发现,直接翻译英文基准无法充分反映文化根植性危害,多语言模型更易遭受跨语言越狱攻击、语码转换攻击及低资源语言中的安全性能退化。这些缺陷源于多语言预训练覆盖不均、本地语言偏好数据不足、安全表征迁移差以及缺乏文化敏感的评估框架。尤其指出,许多低资源语言(尤其是非洲语言)的安全基准数量远低于其他多语言区域。整体揭示持续存在的多语言安全差距,建议未来需建立文化根基的基准、参与式数据收集、均衡的多语言预训练及可扩展的多语言对齐方法。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and multilingual settings than in high-resource languages. In this paper, we conduct a Systematic Literature Review (SLR) of LLM safety alignment in low-resource languages by adopting the PRISMA 2020 methodology. Out of roughly 1,500 papers identified from Semantic Scholar, arXiv, and OpenAlex, 50 relevant studies have been selected and analyzed. Our review is organized around four themes: safety alignment methods, multilingual safety risks, evaluation benchmarks, and cross-lingual transferability. We further propose a taxonomy of safety alignment approaches based on three adaptation mechanisms: data adaptation, objective optimization, and mechanistic alignment. Across literature, translated English benchmarks fail to sufficiently represent culturally rooted harms, and multilingual models are more vulnerable to cross-lingual jailbreaks, code-switching attacks, and safety degradation in underrepresented languages. These failures are driven by several key factors, including uneven multilingual pre-training coverage, insufficient native-language preference data, poor transfer of safety representations, and a lack of culturally aware evaluation frameworks. The review also notes that many low-resource languages, especially African languages, have fewer safety benchmarks available than other multilingual regions. Overall, the results reveal a persistent multilingual safety gap, and suggest that future progress will require culturally grounded benchmarks, participatory data collection, balanced multilingual pre-training, and scalable multilingual alignment methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。