用物理模型提升仿真实体机器人控制迁移效率,省去复杂调参。
Towards bridging the gap: Systematic sim-to-real transfer for diverse legged robots
- 结合强化学习与电机物理能耗模型,设计简洁四分量奖励函数。
- 在13台不同腿足机器人上实现稳定迁移,ANYmal能耗降低32%。
- 无需随机化参数,适合工业级部署和多机型通用控制开发。
腿足机器人需兼具鲁棒运动与能效表现才能应用于真实环境。然而,仿真中训练的控制器常无法可靠迁移到现实,且多数方法忽略执行器特有的能量损耗或依赖复杂的手工设计奖励函数。本文提出一种融合模拟到现实强化学习与永磁同步电机物理能耗模型的框架。该框架仅需少量参数即可刻画仿真与现实间的差异,并采用基于物理原理的四项式奖励函数,平衡电学与机械耗散。通过自下而上的动力学参数辨识研究,覆盖执行器、全机空中轨迹及地面行走。在三类主平台及十台附加机器人上验证,无需动态参数随机化即实现可靠策略迁移。相比现有方法,本方法显著提升能效,在ANYmal上将完整运输成本(Cost of Transport)降低32%(降至1.27)。所有代码、模型与数据集均已公开。
原文摘要 · Abstract (English)
Legged robots must achieve both robust locomotion and energy efficiency to be practical in real-world environments. Yet controllers trained in simulation often fail to transfer reliably, and most existing approaches neglect actuator-specific energy losses or depend on complex, hand-tuned reward formulations. We propose a framework that integrates sim-to-real reinforcement learning with a physics-grounded energy model for permanent magnet synchronous motors. The framework requires a minimal parameter set to capture the simulation-to-reality gap and employs a compact four-term reward with a first-principle-based energetic loss formulation that balances electrical and mechanical dissipation. We evaluate and validate the approach through a bottom-up dynamic parameter identification study, spanning actuators, full-robot in-air trajectories and on-ground locomotion. The framework is tested on three primary platforms and deployed on ten additional robots, demonstrating reliable policy transfer without randomization of dynamic parameters. Our method improves energetic efficiency over state-of-the-art methods, achieving a 32 percent reduction in the full Cost of Transport of ANYmal (value 1.27). All code, models, and datasets are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。