arXiv:2505.12443cs.RO2025-05被引 1

首次系统性攻击多模态大模型导航,90%成功率

BadNAVer: Exploring Jailbreak Attacks On Vision-and-Language Navigation

  • 设计三层攻击框架,构造四类恶意指令组合
  • 在Matterport3D中平均攻击成功率超90%
  • 实机验证可触发机器人物理危害行为

多模态大语言模型(MLLM)在视觉-语言导航(VLN)任务中表现出强大的泛化与推理能力,推动了基于MLLM的导航器发展。然而,这类模型易受越狱攻击:精心设计的提示可绕过安全机制,引发不当输出。在具身场景中,此类漏洞风险更高——不同于文本模型仅生成有害内容,具身智能体可能将恶意指令视为可执行命令,导致真实世界伤害。本文提出首个针对MLLM驱动导航器的系统性越狱攻击范式,构建包含四类意图的恶意查询,并与标准导航指令拼接。在Matterport3D仿真环境中,对五种基于MLLM的导航代理进行评估,平均攻击成功率达90%以上。为验证现实可行性,我们在实体机器人上复现攻击,结果表明即使精心设计的提示也能诱导模型产生有害行为和意图,风险已超越文本污染,可能造成物理伤害。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have recently gained attention for their generalization and reasoning capabilities in Vision-and-Language Navigation (VLN) tasks, leading to the rise of MLLM-driven navigators. However, MLLMs are vulnerable to jailbreak attacks, where crafted prompts bypass safety mechanisms and trigger undesired outputs. In embodied scenarios, such vulnerabilities pose greater risks: unlike plain text models that generate toxic content, embodied agents may interpret malicious instructions as executable commands, potentially leading to real-world harm. In this paper, we present the first systematic jailbreak attack paradigm targeting MLLM-driven navigator. We propose a three-tiered attack framework and construct malicious queries across four intent categories, concatenated with standard navigation instructions. In the Matterport3D simulator, we evaluate navigation agents powered by five MLLMs and report an average attack success rate over 90%. To test real-world feasibility, we replicate the attack on a physical robot. Our results show that even well-crafted prompts can induce harmful actions and intents in MLLMs, posing risks beyond toxic output and potentially leading to physical harm.

越狱攻击具身智能多模态模型导航安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。