用简化模型指导强化学习,让机器人走路无需真人示范且更稳定。
Reduced-Order Model-Guided Reinforcement Learning for Demonstration-Free Humanoid Locomotion
- 先用4自由度简化模型生成节能步态模板
- 再用对抗判别器使真实机器人步态匹配模板分布
- 适合想免真人演示做机器人行走的开发者
我们提出一种两阶段强化学习框架——简化模型引导的强化学习(ROM-GRL),用于人体机器人行走,无需动作捕捉数据或复杂奖励设计。第一阶段通过近端策略优化训练一个4自由度的简化模型(ROM),生成节能步态模板;第二阶段利用这些动态一致的轨迹,指导全身体型策略训练,采用软演员-评论家算法并引入对抗判别器,确保学生模型的五维步态特征分布与ROM演示一致。在1米/秒和4米/秒速度下实验显示,ROM-GRL生成的步态更稳定、对称,跟踪误差显著低于纯奖励基线。通过将轻量级ROM引导信息注入高维策略,该方法弥合了仅靠奖励与基于模仿学习之间的差距,实现了无需人类示范的多样化、自然的人体机器人行为。
原文摘要 · Abstract (English)
We introduce Reduced-Order Model-Guided Reinforcement Learning (ROM-GRL), a two-stage reinforcement learning framework for humanoid walking that requires no motion capture data or elaborate reward shaping. In the first stage, a compact 4-DOF (four-degree-of-freedom) reduced-order model (ROM) is trained via Proximal Policy Optimization. This generates energy-efficient gait templates. In the second stage, those dynamically consistent trajectories guide a full-body policy trained with Soft Actor--Critic augmented by an adversarial discriminator, ensuring the student's five-dimensional gait feature distribution matches the ROM's demonstrations. Experiments at 1 meter-per-second and 4 meter-per-second show that ROM-GRL produces stable, symmetric gaits with substantially lower tracking error than a pure-reward baseline. By distilling lightweight ROM guidance into high-dimensional policies, ROM-GRL bridges the gap between reward-only and imitation-based locomotion methods, enabling versatile, naturalistic humanoid behaviors without any human demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。