arXiv:2411.01349cs.ROcs.LG2024-11被引 4

用域随机生成多样数据,提升扩散策略在人形机器人全身控制中的训练效果

The Role of Domain Randomization in Training Diffusion Policies for Whole-Body Humanoid Control

  • 通过域随机生成不同分布的合成演示数据
  • 人形机器人行走需更大更丰富的数据集才能稳定训练
  • 适合研究机器人运动控制与数据增强的从业者

人形机器人因其与人体结构相似,可利用远程操控、动作捕捉或人类视频等丰富示范数据。然而,从示范中提炼策略仍是难题。尽管扩散策略(DPs)在机器人操作任务中表现优异,其在运动与人形控制中的应用仍不充分。本文在模拟的IsaacGym环境中,通过在不同域随机条件下训练对抗性运动先验(AMP)代理,生成合成示范数据,并对比不同规模与多样性的数据集对DP性能的影响。结果表明,尽管DPs能实现稳定行走,但要成功训练出运动策略,所需数据集比操作任务大得多且需更高多样性,即使在简单场景下亦如此。

原文摘要 · Abstract (English)

Humanoids have the potential to be the ideal embodiment in environments designed for humans. Thanks to the structural similarity to the human body, they benefit from rich sources of demonstration data, e.g., collected via teleoperation, motion capture, or even using videos of humans performing tasks. However, distilling a policy from demonstrations is still a challenging problem. While Diffusion Policies (DPs) have shown impressive results in robotic manipulation, their applicability to locomotion and humanoid control remains underexplored. In this paper, we investigate how dataset diversity and size affect the performance of DPs for humanoid whole-body control. In a simulated IsaacGym environment, we generate synthetic demonstrations by training Adversarial Motion Prior (AMP) agents under various Domain Randomization (DR) conditions, and we compare DPs fitted to datasets of different size and diversity. Our findings show that, although DPs can achieve stable walking behavior, successful training of locomotion policies requires significantly larger and more diverse datasets compared to manipulation tasks, even in simple scenarios.

扩散模型人形机器人域随机化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。