arXiv:2509.06296cs.ROcs.AI2025-09中稿 · publication in IEE…被引 3

用合成数据提升四足机器人训练效率,少跑1900万步就完成训练。

Learning to Walk With Less: A Dyna-Style Approach to Quadrupedal Locomotion

  • 用学习的模型生成短时合成轨迹,结合物理仿真增强训练数据
  • 仅需1964万步模拟即可收敛,比原方法少789万步
  • 适合希望降低仿真成本的机器人控制研究者

传统基于策略的强化学习控制器在四足机器人运动控制中常因数据效率低而需要数百万次与仿真环境的交互才能实现稳定控制。本文提出一种类Dyna的模型增强框架,将基于学习的转移模型生成的短时合成轨迹融入PPO的采样过程,以提升样本效率。合成轨迹以物理仿真为锚点,保障运动稳定性;采用预设调度策略逐步引入合成数据,避免早期训练阶段因模型精度不足导致偏差。通过大量消融实验分析不同数据参数对PPO学习行为的影响。最终在Unitree Go1机器人上验证:仅需1964万步模拟即可收敛(原方法为2753万步),壁时训练时间减少12.24%,且未牺牲策略性能或收敛性。跨平台实验在ANYmal和Unitree Go2上也验证了该框架在高维运动控制中显著减少仿真经验的能力,尽管复杂形态下存在奖励权衡。

原文摘要 · Abstract (English)

Traditional on-policy reinforcement learning (RL) controllers for quadrupedal locomotion often suffer from low data efficiency, requiring millions of interactions with simulated environments to achieve stable control. We integrate model-based techniques that improve sample efficiency by augmenting PPO rollouts with synthetic data in a Dyna-style framework. Our method employs a learned transition model to generate short-horizon synthetic tails for each trajectory, anchored by physics-based simulation to preserve stability. A predefined scheduling strategy gradually integrates synthetic transitions, preventing model usage during early training stages when prediction accuracy is low. Through extensive ablation studies, we analyze how varying data parameters influence PPO's learning behavior. Finally, we validate our method in simulation on a Unitree Go1 robot, reaching convergence with substantially fewer simulation steps (19.64M vs. 27.53M) and a 12.24% reduction in wall-clock training time, without compromising policy performance or convergence. Cross-platform experiments on ANYmal and Unitree Go2 further confirm the framework's ability to learn high-dimensional locomotion control with substantially reduced simulation experience, despite reward trade-offs on complex morphologies.

强化学习四足机器人数据效率模型增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。