发现大模型拒答方向跨语言通用,无需微调即可攻破多语种安全机制。
Refusal Direction is Universal Across Safety-Aligned Languages
- 用14种语言验证拒答方向可跨语言迁移,英文向量能精准绕过其他语言拒答。
- 任意安全对齐语言的拒答向量都能无缝转移至其他语言,成功率接近100%。
- 适合关注多语种安全漏洞与防御机制的研究者阅读。
大型语言模型(LLMs)中的拒答机制对保障安全性至关重要。近期研究发现,拒答行为可通过激活空间中的单一方向实现干预,从而绕过拒答。尽管这一现象主要在英语语境下被验证,但各语言的适当拒答行为同样重要,却仍缺乏理解。本文利用PolyRefuse——一个将恶意与良性英文提示翻译成14种语言构建的多语言安全数据集——研究了跨语言拒答行为。我们发现拒答方向具有惊人的跨语言通用性:从英文中提取的向量可在其他语言中近乎完美地绕过拒答,且无需额外微调。更令人惊讶的是,任意安全对齐语言的拒答方向均可无缝迁移到其他语言。我们将其归因于嵌入空间中拒答向量的平行性,并揭示了跨语言越狱的底层机制。这些发现为构建更鲁棒的多语言安全防御提供了实践指导,并推动了对大型语言模型跨语言漏洞的机制理解。
原文摘要 · Abstract (English)
Refusal mechanisms in large language models (LLMs) are essential for ensuring safety. Recent research has revealed that refusal behavior can be mediated by a single direction in activation space, enabling targeted interventions to bypass refusals. While this is primarily demonstrated in an English-centric context, appropriate refusal behavior is important for any language, but poorly understood. In this paper, we investigate the refusal behavior in LLMs across 14 languages using PolyRefuse, a multilingual safety dataset created by translating malicious and benign English prompts into these languages. We uncover the surprising cross-lingual universality of the refusal direction: a vector extracted from English can bypass refusals in other languages with near-perfect effectiveness, without any additional fine-tuning. Even more remarkably, refusal directions derived from any safety-aligned language transfer seamlessly to others. We attribute this transferability to the parallelism of refusal vectors across languages in the embedding space and identify the underlying mechanism behind cross-lingual jailbreaks. These findings provide actionable insights for building more robust multilingual safety defenses and pave the way for a deeper mechanistic understanding of cross-lingual vulnerabilities in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。