不同语言会改变大模型的越狱漏洞,视觉攻击在西语中更有效。
Same Model, Different Weakness: How Language and Modality Reshape the Jailbreak Attack Surface in Frontier MLLMs

- 对比英西双语下模型越狱攻击效果,发现语言影响安全漏洞分布
- 西语中角色扮演攻击失效,但视觉提示攻击成功率上升
- 安全评估需考虑语言与模态交互,不能分开处理
多模态大语言模型(MLLM)的安全攻击面受语言影响,揭示对齐失败的机制结构。我们首次系统开展跨语言、多模态红队测试,比较美国英语(en-US)与墨西哥西班牙语(es-MX)下四个前沿模型(Claude Sonnet 4.5、GPT-5、Pixtral Large、Qwen Omni)的越狱漏洞。采用固定对抗基准(363种多样提示场景),在纯文本与多模态条件下,由每语言组9名母语标注员共收集52,272条危害评分和二元攻击成功判断。核心发现:语言不均匀放大漏洞。贝叶斯混合效应分析显示,角色扮演等语言框架攻击在西语中显著失效,而视觉显式多模态攻击效果增强,直接指向提示-语言接口而非标注者宽松。这种分离表明语言与视觉对齐失败机制不同,仅换语言即可暴露其差异。实际后果是安全排名无法跨语言保持一致:在es-MX组中,Qwen Omni超过Pixtral Large成为最易攻破模型,此排名反转无法通过英语得分修正恢复;且各代模型绝对攻击成功率虽下降,但差距未缩小。研究证明,将语言与模态视为独立维度的安全评估框架,根本性误判了全球部署的MLLM攻击面,必须重新设计。
原文摘要 · Abstract (English)
The attack surface of a multimodal large language model (MLLM) is language-dependent in ways that reveal the mechanistic structure of alignment failures. We present the first systematic cross-lingual, multimodal red-teaming study comparing jailbreak vulnerability in US English (en-US) and Mexican Spanish (es-MX) across four frontier MLLMs: Claude Sonnet 4.5, GPT-5, Pixtral Large, and Qwen Omni. Using a fixed adversarial benchmark of 363 diverse prompt scenarios administered in text-only and multimodal conditions, we collected 52,272 harm ratings and binary attack success judgements from matched panels of nine native-speaker annotators per language group. Our central finding is that language does not scale vulnerability uniformly. Bayesian mixed-effects analyses reveal that linguistic framing attacks such as role-play become substantially less effective under Spanish prompting, while visually explicit multimodal attacks become more effective, which directly implicates the prompt-language interface rather than global annotator leniency. This dissociation indicates that linguistic and visual alignment failures operate through distinct mechanisms, and that switching language is sufficient to expose that separation. The practical consequence is that safety rankings are not preserved across languages. Qwen Omni overtakes Pixtral Large as the most vulnerable model among es-MX participants, a rank reversal no scalar correction of English-condition scores could recover, and absolute attack success rates have declined across model generations without closing the gaps between them. These findings demonstrate that safety evaluation frameworks treating language and modality as independent dimensions fundamentally misspecify the attack surface of globally deployed MLLMs, and must be redesigned accordingly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。