arXiv:2606.00090cs.ROcs.AI2026-06综述

物理AI系统在执行动作时可能无声失效,本文提出运行时授权框架以保障安全。

Silent Failures in Physical AI: A Literature Review of Runtime Action Authorization for Autonomous Systems

  • 构建运行时授权边界,防止黑箱模型发出危险动作
  • 识别传感器漂移、状态估计误差等导致的无声故障
  • 适合关注机器人安全与可信决策的开发者和研究者

物理AI系统越来越多地将多模态观测、语言指令和学习的世界表征转化为具有物理后果的动作。机器人基础模型、视觉-语言-动作模型及基于世界模型的自主系统可决定车辆、机器人、无人机和工业设备的行为。这一转变暴露出传统AI内容审核与经典机器人安全无法涵盖的安全问题:黑箱模型可能在看似自信、合理且语义一致的情况下发出物理上具有后果的动作,而这种失败可能是无声的,源于传感器漂移、遮挡、状态估计误差、分布偏移、幻觉可操作性或无效的物理假设,直到下游硬件控制器检测到违规才被发现。在具身基础模型、世界模型、机器人仿真、具身安全基准、安全控制、运行时保证、不确定性估计、验证和护栏评估等领域,模型能力与安全机制的发展沿独立路径推进。本文综述中反复出现的缺口是:现有任一研究流均未提供从黑箱物理AI模型到物理执行的完整运行时授权边界。由此分析提出了一个受限问题形式化、沉默物理动作失败的定义、运行时护栏功能的分类法以及评估护栏作为物理AI保障机制的要求。

原文摘要 · Abstract (English)

Physical AI systems increasingly map multimodal observations, language instructions, and learned world representations into physically consequential actions. Robotics foundation models, vision-language-action models, and world-model-based autonomous systems can condition decisions that move vehicles, robots, drones, and industrial machines. This transition exposes a safety problem that is not fully captured by conventional AI content moderation or by classical robot safety alone: a black-box model may issue a physically consequential action while appearing confident, plausible, and semantically aligned. The resulting failure can be silent, arising from sensor drift, occlusion, state-estimation error, distribution shift, hallucinated affordances, or invalid physical assumptions before downstream hardware controllers detect a violation. Across embodied foundation models, world models, robotics simulation, embodied safety benchmarks, safe control, runtime assurance, uncertainty estimation, verification, and guardrail evaluation, model capability and safety mechanisms have advanced along largely separate technical tracks. A recurring gap synthesized here is that no single stream surveyed in this review supplies a complete runtime authorization boundary between black-box Physical AI models and physical execution. The resulting analysis develops a bounded problem formulation, a definition of silent physical-action failure, a taxonomy of runtime guardrail functions, and evaluation requirements for comparing guardrails as Physical AI assurance mechanisms.

物理AI运行时安全机器人安全护栏机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。