arXiv:2604.00202cs.RO2026-04被引 1

用可训练的扩散模型生成机器人动作先验,提升人形机器人自主操作能力。

DreamControl-v2: Simpler and Scalable Autonomous Humanoid Skills via Trainable Guided Diffusion Priors

论文配图:DreamControl-v2: Simpler and Scalable Autonomous Humanoid Skills via Trainable Guided Diffusion Priors
图 1 · 摘自论文原文
  • 在机器人运动空间中训练扩散模型,融合多源数据构建统一动作先验
  • 生成更多样化动作轨迹,实现在模拟与真实机器人上的稳定性能
  • 无需人工筛选,适合希望快速部署复杂操作技能的研究者

开发鲁棒的自主人形机器人运动-操作技能仍是机器人领域未解难题。尽管强化学习在腿式运动中表现成功,但在需要长期规划的复杂交互操作任务中仍面临挑战。近期的DreamControl方法利用现成的人类动作扩散模型作为生成先验,指导强化学习策略训练。本文研究了该动作先验的影响,并提出改进框架:在人形机器人运动空间中直接训练可微分的引导扩散模型,将多样化的真人与机器人数据集融合至统一具身空间。实验表明,该方法因训练数据量更大而捕获更广泛的技能,并实现更自动化的流程,无需人工过滤干预。同时,大规模生成参考轨迹对获得稳健下游强化学习策略至关重要。我们在仿真环境及真实Unitree-G1机器人上进行了充分验证。

原文摘要 · Abstract (English)

Developing robust autonomous loco-manipulation skills for humanoids remains an open problem in robotics. While RL has been applied successfully to legged locomotion, applying it to complex, interaction-rich manipulation tasks is harder given long-horizon planning challenges for manipulation. A recent approach along these lines is DreamControl, which addresses these issues by leveraging off-the-shelf human motion diffusion models as a generative prior to guide RL policies during training. In this paper, we investigate the impact of DreamControl's motion prior and propose an improved framework that trains a guided diffusion model directly in the humanoid robot's motion space, aggregating diverse human and robot datasets into a unified embodiment space. We demonstrate that our approach captures a wider range of skills due to the larger training data mixture and establishes a more automated pipeline by removing the need for manual filtering interventions. Furthermore, we show that scaling the generation of reference trajectories is important for achieving robust downstream RL policies. We validate our approach through extensive experiments in simulation and on a real Unitree-G1.

人形机器人扩散模型强化学习动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。