arXiv:2605.02900cs.CRcs.AI2026-05综述被引 4

系统梳理具身AI在感知到交互全链路的安全风险与防御方法。

Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses

论文配图:Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses
图 1 · 摘自论文原文
  • 构建多层级安全分类框架,覆盖感知到交互的完整链条。
  • 分析500+论文揭示多模态感知融合脆弱性等关键问题。
  • 适合关注机器人安全、可信AI的科研与工程人员参考。

具身人工智能将感知、认知、规划与交互集成于开放世界、高安全要求环境中的智能体。随着其在交通、医疗、工业及辅助机器人等领域的自主化程度提升,保障安全成为技术与社会双重需求。不同于数字AI,具身智能体需应对感知不确定性、知识不完整与动态人机交互,失误可能直接导致物理伤害。本综述系统梳理具身AI安全研究,涵盖从感知、认知到规划、执行与交互的全链条攻击与防御,提出统一多层级分类体系,连接具身特有安全发现与视觉、语言及多模态基础模型的进展。基于超过500篇文献,分析对抗攻击、后门攻击、越狱攻击、硬件级攻击;攻击检测、安全训练与鲁棒推理;以及风险感知的人机交互。研究揭示多个被忽视挑战:多模态感知融合的脆弱性、越狱攻击下规划系统的不稳定性,以及开放场景中人机互信的可靠性。通过建立清晰框架并识别关键研究空白,为构建真正具备能力、自主且安全可靠的现实部署具身智能体提供路线图。

原文摘要 · Abstract (English)

Embodied Artificial Intelligence (Embodied AI) integrates perception, cognition, planning, and interaction into agents that operate in open-world, safety-critical environments. As these systems gain autonomy and enter domains such as transportation, healthcare, and industrial or assistive robotics, ensuring their safety becomes both technically challenging and socially indispensable. Unlike digital AI systems, embodied agents must act under uncertain sensing, incomplete knowledge, and dynamic human-robot interactions, where failures can directly lead to physical harm. This survey provides a comprehensive and structured review of safety research in embodied AI, examining attacks and defenses across the full embodied pipeline, from perception and cognition to planning, action and interaction, and agentic system. We introduce a multi-level taxonomy that unifies fragmented lines of work and connects embodied-specific safety findings with broader advances in vision, language, and multimodal foundation models. Our review synthesizes insights from over 500 papers spanning adversarial, backdoor, jailbreak, and hardware-level attacks; attack detection, safe training and robust inference; and risk-aware human-agent interaction. This analysis reveals several overlooked challenges, including the fragility of multimodal perception fusion, the instability of planning under jailbreak attacks, and the trustworthiness of human-agent interaction in open-ended scenarios. By organizing the field into a coherent framework and identifying critical research gaps, this survey provides a roadmap for building embodied agents that are not only capable and autonomous but also safe, robust, and reliable in real-world deployment.

具身AI安全评测人机交互防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。