攻击者通过摆放带文字的实物物体,操控3D环境中的多模态大模型行为。
Extended to Reality: Prompt Injection in 3D Environments
- 用真实物体摆放实现对3D环境中多模态模型的提示注入攻击。
- 在多种相机路径下,成功诱导多个模型执行恶意指令。
- 揭示现有防御机制对这类物理空间攻击无效,适合安全与机器人研究者参考。
多模态大语言模型(MLLM)已能理解并响应3D环境中的视觉输入,推动了机器人和情境对话代理等应用的发展。当MLLM基于摄像头捕捉的真实世界画面推理时,新的攻击面出现:攻击者可在环境中放置带文字的实体物体,从而覆盖MLLM的原定任务。尽管已有研究关注文本域及数字编辑2D图像中的提示注入,但针对3D环境的研究仍有限。为此,我们提出PI3D,一种通过实体物体摆放而非数字图像修改实现的3D环境提示注入攻击。我们建模并求解一个有效姿态(位置与朝向)问题,使攻击者在保持物体摆放物理合理性的前提下,诱导MLLM执行注入任务。实验表明,PI3D在多种相机轨迹下对多个MLLM均具有效性。我们进一步评估多种防御措施,发现其无法可靠抵御PI3D攻击。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have advanced the capabilities to interpret and act on visual input in 3D environments, empowering diverse applications such as robotics and situated conversational agents. When MLLMs reason over camera-captured views of the physical world, a new attack surface emerges: an attacker can place text-bearing physical objects in the environment to override MLLMs' intended task. While prior work has studied prompt injection in the text domain and through digitally edited 2D images, limited attention has been paid to how these attacks function in 3D environments. To bridge the gap, we introduce PI3D, a prompt injection attack against MLLMs in 3D environments, realized through text-bearing object placement rather than digital image edits. We formulate and solve the problem of identifying an effective pose (position and orientation) for a 3D object with injected text, where the attacker's goal is to induce the MLLM to perform the injected task while ensuring that the object placement remains physically plausible. Experiment results demonstrate that PI3D is an effective attack against multiple MLLMs under diverse camera trajectories. We further evaluate a range of defenses and show that they are not sufficient to reliably defend against PI3D.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。