为机器人引入模块化安全防护,应对大模型带来的开放环境风险。
Modular Safety Guardrails Are Necessary for Foundation-Model-Enabled Robots in the Real World
- 设计监控与干预双层模块化安全框架,覆盖动作、决策与人机协同安全
- 实证表明传统方法在复杂多变场景下无法持续保障安全
- 适合长期部署的智能机器人系统,尤其关注人机共融场景
将基础模型(FMs)融入机器人加速了其在真实世界中的部署,但也带来了由开放语义推理和具身物理行为引发的新安全挑战。这些挑战超越了物理约束满足范畴。本文从三个维度界定FM赋能机器人的安全性:动作安全(物理可行性与约束合规)、决策安全(语义与情境恰当性)、以人为本的安全性(符合人类意图、规范与预期)。我们指出,在任务、环境及人类期望开放、长尾且随时间演化的场景中,现有方法(包括静态验证、单体控制器和端到端学习策略)均不足。为此,提出模块化安全防护机制,包含监控(评估)与干预两层,作为全栈自主系统综合安全的架构基础。此外,强调通过表征对齐与保守性分配实现跨层协同设计的可能性,以实现更快速、更少保守、更有效的安全执行。呼吁社区探索更丰富的防护模块与原则性协同设计策略,推动安全的现实世界物理人工智能部署。
原文摘要 · Abstract (English)
The integration of foundation models (FMs) into robotics has accelerated real-world deployment, while introducing new safety challenges arising from open-ended semantic reasoning and embodied physical action. These challenges require safety notions beyond physical constraint satisfaction. In this paper, we characterize FM-enabled robot safety along three dimensions: action safety (physical feasibility and constraint compliance), decision safety (semantic and contextual appropriateness), and human-centered safety (conformance to human intent, norms, and expectations). We argue that existing approaches, including static verification, monolithic controllers, and end-to-end learned policies, are insufficient in settings where tasks, environments, and human expectations are open-ended, long-tailed, and subject to adaptation over time. To address this gap, we propose modular safety guardrails, consisting of monitoring (evaluation) and intervention layers, as an architectural foundation for comprehensive safety across the autonomy stack. Beyond modularity, we highlight possible cross-layer co-design opportunities through representation alignment and conservatism allocation to enable faster, less conservative, and more effective safety enforcement. We call on the community to explore richer guardrail modules and principled co-design strategies to advance safe real-world physical AI deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。