用一张纸就能劫持机器人,研究物理提示注入攻击的漏洞与防御。
Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots

- 通过四类物理提示攻击,测试机器人视觉语言模型的漏洞。
- 三种主流模型受攻击成功率27%至29%,权威伪装攻击跨模型有效。
- 简单防御措施可提升至100%防护,但可能影响读取场景标签任务。
视觉语言模型(VLM)正被广泛用于机器人系统中作为规划器,将自然语言指令转化为基于视觉理解的可执行动作。这种感知与指令跟随的紧密耦合引入了新攻击面:放置在机器人视野中的对抗性文本可作为间接提示注入到VLM的推理链中。本文对VLM控制的分拣机器人开展系统性物理提示注入攻击研究,提出四类攻击分类——间接标识、任务重定义、权威冒充和冲突注入,并构建包含20个攻击提示的基准,在三种物理场景布局和三种命令表述形式下进行测试,涵盖目的地具体性与规则明确性差异。在三款前沿VLM(GPT-4o、Gemini 2.5 Flash、Qwen3-VL-32B)上共完成5,670次实验,攻击成功率分别为27.0%、29.4%和5.0%。权威冒充和否定类攻击在所有模型间均可迁移。对推理轨迹分析显示,成功攻陷几乎均为显式认知(99.9%承认率),且各模型防御机制不同:Gemini显式拒绝,GPT-4o则依赖感知忽略。评估三种缓解方案:基于提示的防御(75%-100%有效,模型相关)、双阶段验证(85%-100%)和预处理文本掩码(100%)。结果表明,VLM控制的机械操作对人可读的物理标识存在实质性脆弱性,而简单防御能显著降低风险,但需权衡代价。防御措施在本基准中保持通用任务能力,但可能削弱依赖场景标签的任务性能。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are increasingly deployed as planners in robotic systems, where they translate natural-language commands into executable actions grounded in visual scene understanding. This tight coupling between perception and instruction-following introduces a new attack surface: adversarial text placed within the robot's visual field can act as an indirect prompt injection into the VLM's reasoning stack. We present a systematic study of physical prompt injection attacks against VLM-controlled sorting, introducing a four-category taxonomy, indirect signage, task redefinition, authority impersonation, and conflict injection, instantiated as a benchmark of 20 attack prompts evaluated across three physical scene layouts and three command formulations that vary in destination specificity and rule explicitness. Across 5,670 trials on three frontier VLMs (GPT-4o, Gemini 2.5 Flash, Qwen3-VL-32B), attacks succeed at 27.0%, 29.4%, and 5.0% respectively, with authority-impersonating and negation attacks transferring across all three models. Analysis of reasoning traces reveals that successful compromise is nearly always conscious (99.9% acknowledgment rate), and that models defend through structurally different mechanisms, explicit rejection for Gemini, perceptual inattention for GPT-4o. We evaluate three simple mitigations: prompt-based defense (75-100% effective, model-dependent), two-stage verification (85-100%), and pre-processing text masking (100%). Our findings show that VLM-controlled manipulation is meaningfully vulnerable to human-readable physical signage, and that simple defenses substantially reduce risk, though defense choice involves trade-offs. The defenses preserve general task capabilities in our benchmark, but they may impair tasks that require reading in-scene labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。