用数据自举加速训练,让机器人实时规划动态目标避障
Reinforcement Learning with Data Bootstrapping for Dynamic Subgoal Pursuit in Humanoid Robot Navigation
- 高层用强化学习在机器人坐标系选动态子目标
- 低层用模型预测控制生成稳定步态,成功率提升40%以上
- 数据自举生成多样化训练数据,适合复杂环境导航
安全且实时的导航是人形机器人应用的基础。然而,现有双足机器人导航框架往往难以在计算效率与稳定行走所需的精度之间取得平衡。本文提出一种新型分层框架,持续生成动态子目标以引导机器人穿越杂乱环境。该方法包含高层强化学习(RL)规划器,在机器人中心坐标系中选择子目标;以及低层基于模型预测控制(MPC)的规划器,生成稳健步态以抵达这些子目标。为加速并稳定训练过程,引入数据自举技术,利用基于模型的导航方法生成多样且信息丰富的数据集。我们在敏捷机器人公司Digit人形机器人上,通过多个随机障碍物场景的仿真验证了该方法。结果表明,相比原始基于模型的方法及其他学习型方法,本框架显著提升了导航成功率与适应性。
原文摘要 · Abstract (English)
Safe and real-time navigation is fundamental for humanoid robot applications. However, existing bipedal robot navigation frameworks often struggle to balance computational efficiency with the precision required for stable locomotion. We propose a novel hierarchical framework that continuously generates dynamic subgoals to guide the robot through cluttered environments. Our method comprises a high-level reinforcement learning (RL) planner for subgoal selection in a robot-centric coordinate system and a low-level Model Predictive Control (MPC) based planner which produces robust walking gaits to reach these subgoals. To expedite and stabilize the training process, we incorporate a data bootstrapping technique that leverages a model-based navigation approach to generate a diverse, informative dataset. We validate our method in simulation using the Agility Robotics Digit humanoid across multiple scenarios with random obstacles. Results show that our framework significantly improves navigation success rates and adaptability compared to both the original model-based method and other learning-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。