通过学习通用运动特性,让四足机器人在仿真与真实环境间高效迁移。
Learning Task-Invariant Properties via Dreamer: Enabling Efficient Policy Transfer for Quadruped Robots
- 在Dreamer模型中加入不变特性学习,自动提取稳定接触与地形间隙等共性特征。
- 仿真到真实迁移平均提升28.1%,真实爬坡任务成功率从10%升至100%。
- 适合需要快速适配新环境的机器人控制研究者使用。
在复杂动态地形上实现四足机器人稳定行走面临挑战,主要源于仿真环境与真实场景之间的差异。传统仿真到现实迁移方法依赖人工特征设计或昂贵的真实世界微调。为此,本文提出DreamTIP框架,在Dreamer世界模型架构中引入任务不变特性学习,以增强仿真到现实的迁移能力。受大语言模型引导,DreamTIP识别并利用接触稳定性、地形间隙等对动态变化鲁棒且跨任务可迁移的特性,将其作为世界模型的辅助预测目标,使策略学习对底层动态变化不敏感的表示。此外,设计了高效的适应策略,采用混合回放缓冲区和正则化约束,快速校准真实动态,有效缓解表示崩溃与灾难性遗忘。在楼梯、攀爬、倾斜、匍匐等复杂地形上的大量实验表明,DreamTIP在仿真与真实环境中均显著优于现有最优基线。在八种不同仿真迁移任务中平均性能提升28.1%。在真实世界的攀爬任务中,基线方法仅达10%成功率,而本方法达到100%成功率。结果表明,将任务不变特性融入Dreamer学习,为实现鲁棒且可迁移的机器人行走提供了新方案。
原文摘要 · Abstract (English)
Achieving quadruped robot locomotion across diverse and dynamic terrains presents significant challenges, primarily due to the discrepancies between simulation environments and real-world conditions. Traditional sim-to-real transfer methods often rely on manual feature design or costly real-world fine-tuning. To address these limitations, this paper proposes the DreamTIP framework, which incorporates Task-Invariant Properties learning within the Dreamer world model architecture to enhance sim-to-real transfer capabilities. Guided by large language models, DreamTIP identifies and leverages Task-Invariant Properties, such as contact stability and terrain clearance, which exhibit robustness to dynamic variations and strong transferability across tasks. These properties are integrated into the world model as auxiliary prediction targets, enabling the policy to learn representations that are insensitive to underlying dynamic changes. Furthermore, an efficient adaptation strategy is designed, employing a mixed replay buffer and regularization constraints to rapidly calibrate to real-world dynamics while effectively mitigating representation collapse and catastrophic forgetting. Extensive experiments on complex terrains, including Stair, Climb, Tilt, and Crawl, demonstrate that DreamTIP significantly outperforms state-of-the-art baselines in both simulated and real-world environments. Our method achieves an average performance improvement of 28.1% across eight distinct simulated transfer tasks. In the real-world Climb task, the baseline method achieved only a 10\ success rate, whereas our method attained a 100% success rate. These results indicate that incorporating Task-Invariant Properties into Dreamer learning offers a novel solution for achieving robust and transferable robot locomotion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。