arXiv:2608.11146cs.CL2026-08

低资源语言中跨语言安全对齐失效,模型无法有效识别有害内容。

The Illusion of Cross-Lingual Safety in Low-Resource Languages

论文配图:The Illusion of Cross-Lingual Safety in Low-Resource Languages
图 1 · 摘自论文原文
  • 构建跨文化安全数据集LoDNA,对比直译与本地化提示。
  • 跨语言拒绝信号衰减至不足10%,安全机制未有效迁移。
  • 适合关注多语言AI安全的从业者与研究者。

大型语言模型(LLMs)的安全对齐主要在英语中发展,假设这些防护措施可泛化到多语言场景。然而该假设尚未充分验证,尤其在低资源语言中存在漏洞。本文针对四种非洲语言(蒂尼语、豪萨语、阿姆哈拉语、斯瓦希里语),使用新构建的安全数据集LoDNA,对比字面翻译与文化本地化提示。为突破生成式评估局限,提出一种潜在空间几何框架,探测模型隐藏层中的拒绝表征。实验显示,跨语言安全迁移严重受限:多数语言-模型组合中,有害提示的英文拒绝信号保留率低于10%。尽管字面与本地化提示语义高度一致(余弦相似度0.95-0.996),但其表征在各层间发生偏移,表明模型虽编码相关概念,却未传递至安全机制。结果揭示当前多语言安全对齐仅为表面现象,强烈质疑特定低资源语言中存在通用、语言无关的有害性流形的假设。警告:本文包含可能令人不适或有害的示例数据。

原文摘要 · Abstract (English)

Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this assumption remains underexplored and exposes a vulnerability in low-resource languages. We investigate cross-lingual safety transfer in four African languages, Twi, Hausa, Amharic, and Swahili, using LoDNA, a new safety dataset that pairs literal translations with culturally localized prompts. To move beyond generation-based evaluation, we propose a latent geometric framework that probes hidden-state refusal representations in LLMs. Our experimental results show that cross-lingual safety transfer is severely limited; harmful prompts retain less than 10% of the English refusal signal across most language-model pairs. Literal and localized prompts are semantically aligned (cosine 0.95-0.996) but drift across layers, suggesting models encode the concepts without routing them to safety mechanisms. These findings demonstrate that current multilingual safety alignment is superficial, providing strong evidence against the assumption of a universal, language-agnostic harm manifold within the specific low-resource languages studied. Warning: This paper contains example data that may be offensive or harmful.

多语言安全低资源语言模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。