用教师模型指导学生机器人,提升复杂地形行走稳定性与泛化能力。
Teacher Motion Priors: Enhancing Robot Locomotion over Challenging Terrain
- 通过师生框架分离网络设计,简化策略网络结构。
- 在动态地形上实现更稳定的行走,降低开发成本。
- 适合需要快速部署稳健行走策略的人形机器人研究者。
在复杂地形上实现鲁棒的行走仍面临高维控制与环境不确定性的挑战。本文提出一种基于师生范式的教师先验框架,结合模仿学习与辅助任务学习,提升学习效率与泛化能力。不同于依赖编码器状态嵌入的传统方法,该框架解耦网络设计,简化策略网络与部署流程。首先使用特权信息训练高性能教师策略以获得可泛化的运动技能;再通过生成对抗机制将教师的运动分布迁移至仅依赖噪声本体感知数据的学生策略,缓解分布偏移导致的性能下降。此外,辅助任务学习增强了学生策略的特征表示,加速收敛并提升对不同地形的适应性。该框架在人形机器人上验证,显著提升了动态地形上的行走稳定性,并大幅降低开发成本。本工作为部署鲁棒行走策略提供了实用解决方案。
原文摘要 · Abstract (English)
Achieving robust locomotion on complex terrains remains a challenge due to high dimensional control and environmental uncertainties. This paper introduces a teacher prior framework based on the teacher student paradigm, integrating imitation and auxiliary task learning to improve learning efficiency and generalization. Unlike traditional paradigms that strongly rely on encoder-based state embeddings, our framework decouples the network design, simplifying the policy network and deployment. A high performance teacher policy is first trained using privileged information to acquire generalizable motion skills. The teacher's motion distribution is transferred to the student policy, which relies only on noisy proprioceptive data, via a generative adversarial mechanism to mitigate performance degradation caused by distributional shifts. Additionally, auxiliary task learning enhances the student policy's feature representation, speeding up convergence and improving adaptability to varying terrains. The framework is validated on a humanoid robot, showing a great improvement in locomotion stability on dynamic terrains and significant reductions in development costs. This work provides a practical solution for deploying robust locomotion strategies in humanoid robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。