arXiv:2607.28226cs.CRcs.AI2026-07被引 3

世界模型让具身智能更聪明,但也带来全新安全风险。

Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation

  • 从数据到执行全程追踪世界模型的安全威胁
  • 攻击可伪装成数据污染或指令注入,导致物理行为失控
  • 适合关注AI安全与具身系统研发的从业者

世界模型为具身人工智能提供预测核心:将观测压缩为状态,模拟动作条件下的未来,并实现超越反应式控制的规划。然而这一预测层开辟了新的安全边界——攻击可从数据、传感器、提示或反馈中渗透并影响物理行为。本文不将世界模型视为孤立组件,而是贯穿其全生命周期:从数据构建与表征学习,到状态锚定与想象,再到轨迹评估、执行及通过记忆与工具的长期适应。我们发现,常见攻击类型如投毒、后门、对抗样本、传感器欺骗、提示注入、轨迹篡改和供应链攻击,在污染世界状态、学习动态、可操作性估计或安全成本时呈现出新特征。同时揭示双重性:世界模型可作为运行时安全屏障,但一旦被攻破或过度依赖,会生成虚假的安全幻觉。本文提出生命周期分类法,将现有攻击映射至世界模型安全属性,设计安全失效评估协议,并在来源可信、鲁棒锚定、不确定性感知预测、轨迹过滤、反馈审计和部署保障层面构建防御体系。

原文摘要 · Abstract (English)

World models give embodied AI a predictive core: they compress observations into states, simulate action-conditioned futures, and enable planning beyond reactive control. This predictive layer, however, opens a new security boundary-compromise can propagate from data, sensors, prompts, or feedback into physical action. Rather than treating world models as an isolated component, this survey traces threats across their entire lifecycle-from data construction and representation learning, through state grounding and imagination, to trajectory evaluation, execution, and long-term adaptation via memory and tools. We show that familiar attack families: poisoning, backdoors, adversarial examples, sensor spoofing, prompt injection, trajectory manipulation, and supply-chain attacks take on distinct meanings when they corrupt world states, learned dynamics, affordance estimates, or safety costs. We also highlight a duality: world models can serve as runtime safety shields, yet when compromised or over-trusted they generate predictive safety illusions. The survey offers a lifecycle taxonomy, maps existing attacks to world-model security properties, outlines evaluation protocols for safety failures, and structures defenses across provenance, robust grounding, uncertainty-aware prediction, trajectory gating, feedback auditing, and deployment assurance.

AI安全具身智能世界模型攻击防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。