用扩散模型生成轨迹引导机器人完成长序列操作,减少误差累积。
Diffusion Trajectory-guided Policy for Long-horizon Robot Manipulation
- 用扩散模型生成任务相关的2D操作轨迹,指导策略学习
- 在CALVIN基准上成功率达85%,比顶尖方法高25%
- 无需预训练,适合真实机器人长时序任务
近期视觉-语言-动作模型(VLA)推动了机器人模仿学习的发展,但高昂的数据收集成本和有限的示范数据限制了泛化能力,现有模仿学习方法在分布外场景中表现不佳,尤其在长时序任务中。核心挑战在于如何缓解模仿学习中的误差累积问题,避免长时间轨迹上的级联失败。为此,我们提出扩散轨迹引导策略(DTP)框架,通过扩散模型生成2D轨迹,为长时序任务提供轨迹级引导以减少误差积累。该方法采用两阶段设计:首先训练一个生成式视觉-语言模型以生成基于扩散的轨迹,随后利用这些轨迹微调模仿策略。在CALVIN基准上的实验表明,DTP在从零开始训练的情况下,成功率达到85%,比现有最优基线提升25%。此外,DTP显著提升了真实机器人上的性能。
原文摘要 · Abstract (English)
Recently, Vision-Language-Action models (VLA) have advanced robot imitation learning, but high data collection costs and limited demonstrations hinder generalization and current imitation learning methods struggle in out-of-distribution scenarios, especially for long-horizon tasks. A key challenge is how to mitigate compounding errors in imitation learning, which lead to cascading failures over extended trajectories. To address these challenges, we propose the Diffusion Trajectory-guided Policy (DTP) framework, which generates 2D trajectories through a diffusion model to guide policy learning for long-horizon tasks. By leveraging task-relevant trajectories, DTP provides trajectory-level guidance to reduce error accumulation. Our two-stage approach first trains a generative vision-language model to create diffusion-based trajectories, then refines the imitation policy using them. Experiments on the CALVIN benchmark show that DTP outperforms state-of-the-art baselines by 25% in success rate, starting from scratch without external pretraining. Moreover, DTP significantly improves real-world robot performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。