arXiv:2503.03480cs.ROcs.AI2025-03NeurIPS被引 64

让视觉语言动作模型更安全:通过约束学习显式整合安全规则。

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning

  • 基于约束马尔可夫决策过程,从对抗性风险中优化安全策略。
  • 安全违规累积成本降低83.58%,任务成功率提升3.85%。
  • 适用于机器人高风险场景,适合关注安全对齐的研究者。

视觉语言动作模型(VLAs)在作为通用机器人策略方面展现出潜力,但在真实部署中存在对环境、机器人自身及人类造成伤害的极端安全风险。如何将安全约束显式融入VLAs?我们提出集成安全方法(ISA),系统建模安全需求,主动诱导多样化的不安全行为,通过安全强化学习有效约束VLA策略,并以针对性评估严格验证其安全性。基于约束马尔可夫决策过程(CMDP)框架,ISA从极小极大视角优化对诱发的安全风险。经此综合方法对齐后的策略具备以下关键特性:(I) 有效实现安全与性能的权衡,相比当前最优方法,安全违规累积成本降低83.58%,同时任务成功率提升3.85%;(II) 强大的安全保障能力,能缓解长尾风险并应对极端失效场景;(III) 学习到的安全行为对多种分布外扰动具有鲁棒泛化能力。有效性在长时间跨度移动操作任务上进行验证。数据、模型及新提出的基准环境已公开于https://pku-safevla.github.io。

原文摘要 · Abstract (English)

Vision-language-action models (VLAs) show potential as generalist robot policies. However, these models pose extreme safety challenges during real-world deployment, including the risk of harm to the environment, the robot itself, and humans. How can safety constraints be explicitly integrated into VLAs? We address this by exploring an integrated safety approach (ISA), systematically modeling safety requirements, then actively eliciting diverse unsafe behaviors, effectively constraining VLA policies via safe reinforcement learning, and rigorously assuring their safety through targeted evaluations. Leveraging the constrained Markov decision process (CMDP) paradigm, ISA optimizes VLAs from a min-max perspective against elicited safety risks. Thus, policies aligned through this comprehensive approach achieve the following key features: (I) effective safety-performance trade-offs, reducing the cumulative cost of safety violations by 83.58% compared to the state-of-the-art method, while also maintaining task success rate (+3.85%). (II) strong safety assurance, with the ability to mitigate long-tail risks and handle extreme failure scenarios. (III) robust generalization of learned safety behaviors to various out-of-distribution perturbations. The effectiveness is evaluated on long-horizon mobile manipulation tasks. Our data, models and newly proposed benchmark environment are available at https://pku-safevla.github.io.

机器人安全多模态对齐强化学习约束学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。