通过分层优化约束强化学习,提升足式机器人在复杂环境下的安全行走能力。
Whole-Body Constrained Learning for Legged Locomotion via Hierarchical Optimization
- 采用分层优化框架,将硬约束与软约束融入强化学习
- 在雪坡、台阶等户外环境中实现安全稳定通行
- 兼顾仿真到现实的迁移能力与复杂地形适应性
强化学习(RL)在多种挑战性环境中展现出优异的足式机器人运动性能。然而,由于仿真到现实的差距以及缺乏可解释性,未加约束的RL策略在真实场景中仍存在关节碰撞、扭矩过大或低摩擦环境下足部打滑等安全隐患,限制了其在行星探测、核设施巡检和深海作业等高安全要求任务中的应用。本文设计了一种基于分层优化的全身跟随控制方法,将硬约束与软约束整合进强化学习框架,提升机器人运动的安全性。利用模型预测控制的优势,该方法可在训练或部署阶段定义多种类型约束,支持策略微调,并缓解仿真到现实的迁移难题。同时保持强化学习在复杂非结构化环境中的鲁棒性。所训练的带约束策略已在六足机器人上部署,测试于雪地斜坡、台阶等多种户外场景,验证了方法出色的越障能力和安全性。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has demonstrated impressive performance in legged locomotion over various challenging environments. However, due to the sim-to-real gap and lack of explainability, unconstrained RL policies deployed in the real world still suffer from inevitable safety issues, such as joint collisions, excessive torque, or foot slippage in low-friction environments. These problems limit its usage in missions with strict safety requirements, such as planetary exploration, nuclear facility inspection, and deep-sea operations. In this paper, we design a hierarchical optimization-based whole-body follower, which integrates both hard and soft constraints into RL framework to make the robot move with better safety guarantees. Leveraging the advantages of model-based control, our approach allows for the definition of various types of hard and soft constraints during training or deployment, which allows for policy fine-tuning and mitigates the challenges of sim-to-real transfer. Meanwhile, it preserves the robustness of RL when dealing with locomotion in complex unstructured environments. The trained policy with introduced constraints was deployed in a hexapod robot and tested in various outdoor environments, including snow-covered slopes and stairs, demonstrating the great traversability and safety of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。