多语言大模型安全漏洞暴露,18种语言测试揭示不同攻击方式的失效机制。
Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety

- 构建跨18语言、4资源层级的多语言越狱基准测试
- 发现低资源语言中越狱成功率超90%,且存在安全阈值跃迁
- 适合关注多语言AI安全与模型对齐的研究者
大型语言模型在多语言环境下的安全对齐仍具脆弱性:英文中能被可靠拒绝的提示,在非英语和低资源语境中可能引发有害响应。我们提出 extsc{Minionese},一个覆盖18种语言、4类资源层级和4种扰动类型(标准翻译、代码切换、音译、翻译腔)的多语言越狱基准,并结合几何机制分析各语言层级的拒绝失败模式。结果显示,每种攻击方式呈现独特脆弱性:音译漏洞由文字体系决定,代码切换在最低资源层级仍有效,所有模型在第2与第3资源层级间均出现显著安全跃迁。机理上,低资源越狱通过将有害内容投射至几何错位子空间,使其无法充分激活拒绝方向,导致拒绝机制未触发但依然完整。研究证明仅以英文评估安全不可靠,需考虑语系、扰动类型与逐语言对齐覆盖。基准与分析代码见https://github.com/Brentkong/Minionese-Comprehensive-Benchmark-and-Mechanistic-Study-of-Multilingual-LLM-Safety.git。
原文摘要 · Abstract (English)
Safety alignment in large language models remains brittle across languages: prompts reliably refused in English can elicit harmful compliance in non-English and low-resource settings. We introduce \textsc{Minionese}, a multilingual jailbreak benchmark spanning 18 languages, 4 resource tiers, and 4 perturbation types (standard translation, code-switching, transliteration, and translationese), paired with a geometric mechanistic analysis of refusal failure across language tiers. We show that each attack type produces a distinct vulnerability profile: transliteration vulnerability is mediated by script identity, code-switching maintains effectiveness through the lowest-resource tier, and a sharp safety regime transition between Tiers 2 and 3 is consistent across all models. Mechanistically, low-resource jailbreaks succeed by routing harmful content through a geometrically misaligned subspace that projects insufficiently onto the refusal directions, leaving the refusal mechanism intact but untriggered. These findings show that English-only safety evaluations are insufficient; they require accounting for script family, perturbation type, and per-language alignment coverage. The benchmark and analysis code is at https://github.com/Brentkong/Minionese-Comprehensive-Benchmark-and-Mechanistic-Study-of-Multilingual-LLM-Safety.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。