arXiv:2606.26855cs.RO2026-06被引 1

用扩散模型生成轨迹,让机器人自动学习复杂人形操作技能。

Humanoid-DART: Humanoid Loco-Manipulation using Diffusion-guided Augmentation through Relabeling and Tracking

论文配图:Humanoid-DART: Humanoid Loco-Manipulation using Diffusion-guided Augmentation through Relabeling and Tracking
图 1 · 摘自论文原文
  • 通过扩散模型生成动作轨迹,结合强化学习追踪目标路径。
  • 仅需少量示范即可扩展技能库,实现端到端的自监督学习。
  • 适合研究人形机器人自主操控与少样本学习的学者。

模仿人类示范已成为学习人形机器人运动与操作策略的主流方法。然而,由于收集多样化示范成本高昂,且需持续人工干预修正策略错误,该方法在规模化时面临挑战。本文提出一种自监督框架,从稀疏示范出发,逐步拓展行为能力,实现仅需少量专家指导的目标条件策略学习。该方法融合基于扩散模型的轨迹生成与强化学习,利用后者追踪扩散模型生成的多种运动-操作技能的目标条件轨迹。通过大量消融实验与前沿方法对比,验证了框架在多类人形运动-操作任务上的有效性。

原文摘要 · Abstract (English)

Imitating human demonstrations has emerged as a dominant paradigm for learning humanoid loco-manipulation policies. However, scaling these approaches remains challenging due to the high cost of collecting diverse demonstrations and the need for continual human intervention to correct policy failures. In this paper, we present a self-supervised framework that bootstraps from sparse demonstrations and progressively expands its behavioral repertoire, enabling the learning of a goal-conditioned policy that automatically explores the goal space with minimal expert supervision. Our approach combines diffusion-based trajectory generation with reinforcement learning, where the latter is used to track goal-conditioned trajectories produced by the diffusion model for a range of loco-manipulation skills. Through extensive ablation studies and comparisons with state-of-the-art methods, we demonstrate the effectiveness of our framework on multiple humanoid loco-manipulation skills.

人形机器人扩散模型自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。