arXiv:2606.03793cs.CLcs.CV2026-06

多语言多模态模型存在跨语言攻击漏洞,安全表现依赖真实理解而非识别失败。

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models

论文配图:Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models
图 1 · 摘自论文原文
  • 通过梯度攻击测试12种语言,发现对抗样本具强跨语言迁移性
  • 低资源语言看似安全实为视觉识别失败,称为'安全幻觉'
  • 全程多语言训练的模型(如Qwen3-VL)才具备真正跨语言拒答能力

多模态大语言模型将视觉感知融入语言推理,带来持续的攻击面,现有研究主要聚焦英文任务,忽视多语言行为。本文系统评估12种语言下开源多模态大模型的鲁棒性与多模态安全性,这些模型通过指令微调获得多语言能力。梯度攻击揭示出可迁移的多语言脆弱性:在一种语言中优化的对抗图像仍能引发其他语言的失效,表现出强跨语言迁移性。多模态安全性随模型对有害指令的检索与理解能力而异。当有害意图以文本形式出现时,语言基础更强的语言更常产生滥用响应;而弱语言产生较少不安全输出。当有害内容嵌入图像作为文字时,英语文本被视觉编码器可靠识别并执行,非英语文本则极少被解析。因此低资源语言看似更安全,实为理解与视觉对齐失败所致,我们称之为‘安全-失败’现象。相比之下,全程训练阶段融入多语言能力的模型(如Qwen3-VL)展现出真正的跨语言安全,能在各语言中主动拒绝,而非掩盖理解失败。浅层多语言适配(如翻译指令数据微调)可能制造表面理解,导致低资源语言的虚假安全;而贯穿训练全过程的深层整合才能实现真正的多语言安全对齐。

原文摘要 · Abstract (English)

Multimodal Large Language Models integrate visual perception into language reasoning, introducing a continuous attack surface susceptible to adversarial attacks. Prior work on MLLM robustness has focused largely on English-centric tasks, leaving multilingual behaviour unexplored. We address this gap through a systematic study of adversarial robustness and multimodal safety across 12 diverse languages, evaluating open-source MLLMs that acquire multilingual capability through instruction tuning. Gradient-based attacks reveal a transferable multilingual vulnerability: adversarial images optimized in one language continue to induce failure in others, demonstrating strong cross-lingual transferability. Multilingual safety further varies with how effectively a model retrieves or interprets harmful instructions. When harmful intent is issued through text, languages with stronger linguistic grounding more often elicit misuse-enabling responses, while weaker languages produce fewer unsafe outputs. When embedded in the image as typographic content, English scripts are reliably recognised and followed, whereas non-English scripts are rarely parsed by the vision encoder. Lower-resource languages may therefore appear safer, but this is an artefact of comprehension and visual-grounding failures rather than genuine alignment, a phenomenon we term safety-by-failure. In contrast, MLLMs that build multilingual capability throughout their training stages rather than only at instruction tuning, such as Qwen3-VL, exhibit genuine cross-lingual safety, maintaining active refusal across languages rather than masking comprehension failure. Shallow multilingual adaptation, such as fine-tuning on translated instruction data, may produce surface-level understanding that creates illusory safety in low-resource languages; deeper integration across training stages leads to genuine multilingual safety alignment.

多模态安全对齐跨语言对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。