自动挑选高难度地形训练机器人,提升越障效率与真实世界适应力。
VertiSelector: Automatic Curriculum Learning for Wheeled Mobility on Vertically Challenging Terrain
- 根据学习误差动态选择高挑战性地形进行训练
- 在仿真中实现23.08%成功率提升,泛化能力更强
- 适合需要高效强化学习的移动机器人研究者
强化学习(RL)有望通过端到端试错学习实现极端非结构化地形下的轮式移动,无需复杂的动力学建模与规划。然而,现有方法在大量手动设计的仿真环境中训练时样本效率低,且难以泛化到现实世界。为此,我们提出VertiSelector(VS),一种自动课程学习框架,通过有选择地采样训练地形来提升学习效率与泛化能力。VS在重访时优先选择具有更高时间差分(TD)误差的垂直挑战性地形,使机器人在能力边界持续学习。通过动态调整采样重点,VS显著提升了在基于Chrono多物理引擎构建的VW-Chrono仿真器中的样本效率与泛化性能。此外,我们在Verti-4-Wheeler平台上进行了仿真与实物实验,结果表明,使用VS可使成功率达23.08%的提升,并在真实环境中保持鲁棒性。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) has the potential to enable extreme off-road mobility by circumventing complex kinodynamic modeling, planning, and control by simulated end-to-end trial-and-error learning experiences. However, most RL methods are sample-inefficient when training in a large amount of manually designed simulation environments and struggle at generalizing to the real world. To address these issues, we introduce VertiSelector (VS), an automatic curriculum learning framework designed to enhance learning efficiency and generalization by selectively sampling training terrain. VS prioritizes vertically challenging terrain with higher Temporal Difference (TD) errors when revisited, thereby allowing robots to learn at the edge of their evolving capabilities. By dynamically adjusting the sampling focus, VS significantly boosts sample efficiency and generalization within the VW-Chrono simulator built on the Chrono multi-physics engine. Furthermore, we provide simulation and physical results using VS on a Verti-4-Wheeler platform. These results demonstrate that VS can achieve 23.08% improvement in terms of success rate by efficiently sampling during training and robustly generalizing to the real world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。