系统梳理具身智能体安全风险,从信任边界出发划分攻击面与防御重点。
Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation
- 按信任边界划分五层十二类攻击面,突破传统机制分类局限。
- 58起攻击记录显示多模态感知和动作接口最易受攻,防御多集中于运行时保护。
- 适合研究具身智能体安全的学者、工程师及评估者参考,尤其关注长期记忆与多机协同漏洞。
基础模型越来越多地用于具身智能体的感知、推理、规划与行为生成,带来从数字输入到物理行为的安全风险传播。现有综述多按逃逸、提示注入、后门、污染、对抗样本等机制分类,但未能一致识别攻击首次渗入具身控制环的位置。本文提出以信任边界为中心的综述框架,基于首个被攻破的信任边界原则,将攻击面与攻击机制分离,构建涵盖模型供应链、用户指令、上下文与记忆、物理语义环境、多模态感知、世界状态、内部推理、任务规划、动作接口、中间件、多智能体通信和执行控制的五层十二类攻击面。基于2026年8月15日前收集的58起攻击记录与61项防御记录,分析代表性攻击、跨层传播、防御部署位置与评估实践。定量分析表明,攻击研究集中于多模态感知与动作接口,而防御则高度集中于动作层与运行时保护;上下文与长期记忆、中间件与网络、世界状态完整性、多智能体信任等领域仍较薄弱。最后提出状态溯源、组合式防御、长时程攻击传播、物理可实现性、拜占庭式多机器人行为与闭环统一评估等开放挑战。
原文摘要 · Abstract (English)
Foundation models are increasingly used for perception, reasoning, planning, and action generation in embodied agents, creating security risks that can propagate from digital inputs to physical behavior. Existing surveys often organize threats by mechanisms such as jailbreaks, prompt injection, backdoors, poisoning, or adversarial examples, but these categories do not consistently identify where an adversary first enters the embodied control loop. We present a trust-boundary-centric survey of foundation-model-powered embodied-agent security. Using a first-compromised-trust-boundary principle, we separate attack surface from attack mechanism and organize the system into five layers and twelve attack surfaces spanning the model supply chain, user instructions, context and memory, physical semantic environments, multimodal perception, world state, internal reasoning, task planning, action interfaces, middleware, multi-agent communication, and execution control. Based on 58 attack records and 61 defense records collected through August 15, 2026, we analyze representative attacks, cross-layer propagation, defense placement, and evaluation practices. Our quantitative analysis shows that attack research is concentrated on multimodal perception and action interfaces, while defenses are especially concentrated on action-level and runtime protection. Context and long-term memory, middleware and networking, world-state integrity, and multi-agent trust remain comparatively underexplored. We conclude with open challenges in state provenance, compositional defenses, long-horizon attack propagation, physical realizability, Byzantine multi-robot behavior, and unified closed-loop evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。