让四足机器人在野外实时学习并安全行走,靠自适应控制与安全备份协同实现。
Runtime Learning of Quadruped Robots in Wild Environments
- 用深度强化学习自学动作策略,同时由物理模型控制器监督安全。
- 实测表明在复杂环境中比现有安全强化学习方法更稳定高效。
- 适合需要自主适应野外环境的机器人研发人员参考。
本文提出一种四足机器人在动态野外环境中的运行时学习框架,使机器人能安全地自主学习与适应。该框架集成感知、导航与控制,形成闭环系统。其核心创新在于控制模块中两个交互互补的组件:高性能(HP)-学生与高保障(HA)-教师。HP-学生是一个深度强化学习(DRL)代理,通过自我学习与教学式学习构建安全且高性能的动作策略;HA-教师是基于简化可验证物理模型的控制器,负责向HP-学生传授安全性知识,并在紧急情况下提供安全运动的后备保障。HA-教师具备实时物理模型、实时动作策略与实时控制目标,能有效响应真实野外环境变化,确保系统安全。框架还包含协调器,以高效管理两者的协作。实验使用Unitree Go2机器人在Nvidia Isaac Gym中进行,对比当前最先进的安全DRL方法,验证了所提框架的有效性。
原文摘要 · Abstract (English)
This paper presents a runtime learning framework for quadruped robots, enabling them to learn and adapt safely in dynamic wild environments. The framework integrates sensing, navigation, and control, forming a closed-loop system for the robot. The core novelty of this framework lies in two interactive and complementary components within the control module: the high-performance (HP)-Student and the high-assurance (HA)-Teacher. HP-Student is a deep reinforcement learning (DRL) agent that engages in self-learning and teaching-to-learn to develop a safe and high-performance action policy. HA-Teacher is a simplified yet verifiable physics-model-based controller, with the role of teaching HP-Student about safety while providing a backup for the robot's safe locomotion. HA-Teacher is innovative due to its real-time physics model, real-time action policy, and real-time control goals, all tailored to respond effectively to real-time wild environments, ensuring safety. The framework also includes a coordinator who effectively manages the interaction between HP-Student and HA-Teacher. Experiments involving a Unitree Go2 robot in Nvidia Isaac Gym and comparisons with state-of-the-art safe DRLs demonstrate the effectiveness of the proposed runtime learning framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。