arXiv:2506.11470cs.RO2025-06被引 10

用扩散模型+强化学习,让不同腿型机器人学会通用行走

Multi-Loco: Unifying Multi-Embodiment Legged Locomotion via Reinforcement Learning Augmented Diffusion

  • 用生成式扩散模型捕捉跨机器人共有的运动模式
  • 残差策略提升动作精度,实测平均性能提升10.35%
  • 适合需要多机器人通用控制的科研与工程场景

由于观测/动作维度和系统动力学差异,跨多种腿型机器人泛化运动策略是一大挑战。本文提出 Multi-Loco,一种结合形态无关生成扩散模型与轻量级强化学习(RL)残差策略的统一框架。扩散模型从多机器人异构数据中学习形态不变的运动模式,提升泛化性与鲁棒性;共享的残差策略对扩散模型生成的动作进行优化,增强任务感知性能与真实部署适应性。我们在包含四类四足机器人的仿真与真实世界实验中验证方法,相比标准PPO框架,将平均回报提升10.35%,在轮足行走任务中最高达13.57%。结果表明,跨形态数据与复合生成架构有助于学习稳健、可泛化的运动技能。

原文摘要 · Abstract (English)

Generalizing locomotion policies across diverse legged robots with varying morphologies is a key challenge due to differences in observation/action dimensions and system dynamics. In this work, we propose Multi-Loco, a novel unified framework combining a morphology-agnostic generative diffusion model with a lightweight residual policy optimized via reinforcement learning (RL). The diffusion model captures morphology-invariant locomotion patterns from diverse cross-embodiment datasets, improving generalization and robustness. The residual policy is shared across all embodiments and refines the actions generated by the diffusion model, enhancing task-aware performance and robustness for real-world deployment. We evaluated our method with a rich library of four legged robots in both simulation and real-world experiments. Compared to a standard RL framework with PPO, our approach -- replacing the Gaussian policy with a diffusion model and residual term -- achieves a 10.35% average return improvement, with gains up to 13.57% in wheeled-biped locomotion tasks. These results highlight the benefits of cross-embodiment data and composite generative architectures in learning robust, generalized locomotion skills.

强化学习扩散模型机器人控制多机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。