用视觉动作数据生成有安全保证的机器人控制策略
Semi-Supervised Safe Visuomotor Policy Synthesis using Barrier Certificates
- 结合无模型学习与基于模型控制,用屏障证书保证安全
- 仅需部分安全标签即可训练出可证明安全的策略
- 适合需要安全性的自主机器人系统开发
在现代机器人领域,真实场景中状态空间信息不准确的问题促使研究者利用视觉动作观测来提供安全保证。虽然监督学习方法(如模仿学习)能在视觉动作观测基础上合成控制策略,但需要完整数据集的安全标签,且无法提供形式化安全保证。而传统控制理论方法(如控制屏障函数CBFs和哈密顿-雅可比可达性)虽能提供形式化安全保证,却依赖对系统动态的精确知识,难以应用于高维视觉动作数据。为此,本文提出一种新型半监督安全视觉动作策略合成方法,基于屏障证书融合无模型监督学习与基于模型控制的优势。该框架无需完整数据集的安全标签,即可合成可证明安全的控制器,并确保屏障证书与策略的完备性。通过倒立摆系统和自主移动机器人避障两个案例验证了该方法的有效性。
原文摘要 · Abstract (English)
In modern robotics, addressing the lack of accurate state space information in real-world scenarios has led to a significant focus on utilizing visuomotor observation to provide safety assurances. Although supervised learning methods, such as imitation learning, have demonstrated potential in synthesizing control policies based on visuomotor observations, they require ground truth safety labels for the complete dataset and do not provide formal safety assurances. On the other hand, traditional control-theoretic methods like Control Barrier Functions (CBFs) and Hamilton-Jacobi (HJ) Reachability provide formal safety guarantees but depend on accurate knowledge of system dynamics, which is often unavailable for high-dimensional visuomotor data. To overcome these limitations, we propose a novel approach to synthesize a semi-supervised safe visuomotor policy using barrier certificates that integrate the strengths of model-free supervised learning and model-based control methods. This framework synthesizes a provably safe controller without requiring safety labels for the complete dataset and ensures completeness guarantees for both the barrier certificate and the policy. We validate our approach through distinct case studies: an inverted pendulum system and the obstacle avoidance of an autonomous mobile robot.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。