用环境物体做触发器,让视觉语言模型机器人突然执行恶意指令。
BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
- 用对比学习强化触发器识别,提升攻击精准度。
- 在多个任务中攻击成功率最高达80%,且正常任务表现不受影响。
- 适合研究AI安全或智能体防御的开发者关注。
视觉语言模型(VLM)的进步推动了具身智能体的发展,使其能直接从视觉输入中感知、推理并规划任务导向动作。然而,这类视觉驱动的具身智能体引入了新的攻击面:视觉后门攻击,即智能体在无触发时表现正常,一旦场景中出现特定视觉触发器,便持续执行攻击者指定的多步策略。我们提出BEAT,首个利用环境中物体作为触发器向VLM-based具身智能体注入视觉后门的框架。不同于文本触发器,物体触发器因视角和光照变化多样,难以可靠植入。BEAT通过(1)构建涵盖多种场景、任务和触发位置的训练集以暴露触发器变异性,(2)采用两阶段训练:先进行监督微调(SFT),再引入新颖的对比触发学习(CTL)。CTL将触发器判别建模为带触发与无触发输入之间的偏好学习,显式增强决策边界,确保后门精准激活。在多个具身智能体基准和VLM上,BEAT实现最高80%的攻击成功率,同时保持强良性任务性能,并对分布外触发位置具有可靠泛化能力。值得注意的是,相较于朴素SFT,CTL在后门数据有限时可将激活准确率提升最高达39%。这些发现揭示了基于VLM的具身智能体中一个关键且未被探索的安全风险,强调在实际部署前需建立鲁棒防御机制。
原文摘要 · Abstract (English)
Recent advances in Vision-Language Models (VLMs) have propelled embodied agents by enabling direct perception, reasoning, and planning task-oriented actions from visual inputs. However, such vision-driven embodied agents open a new attack surface: visual backdoor attacks, where the agent behaves normally until a visual trigger appears in the scene, then persistently executes an attacker-specified multi-step policy. We introduce BEAT, the first framework to inject such visual backdoors into VLM-based embodied agents using objects in the environments as triggers. Unlike textual triggers, object triggers exhibit wide variation across viewpoints and lighting, making them difficult to implant reliably. BEAT addresses this challenge by (1) constructing a training set that spans diverse scenes, tasks, and trigger placements to expose agents to trigger variability, and (2) introducing a two-stage training scheme that first applies supervised fine-tuning (SFT) and then our novel Contrastive Trigger Learning (CTL). CTL formulates trigger discrimination as preference learning between trigger-present and trigger-free inputs, explicitly sharpening the decision boundaries to ensure precise backdoor activation. Across various embodied agent benchmarks and VLMs, BEAT achieves attack success rates up to 80%, while maintaining strong benign task performance, and generalizes reliably to out-of-distribution trigger placements. Notably, compared to naive SFT, CTL boosts backdoor activation accuracy up to 39% under limited backdoor data. These findings expose a critical yet unexplored security risk in VLM-based embodied agents, underscoring the need for robust defenses before real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。