用视觉语言模型指导扩散策略,提升长程机械臂操作的鲁棒性
VLM-TDP: VLM-guided Trajectory-conditioned Diffusion Policy for Robust Long-Horizon Manipulation
- 通过视觉语言模型拆解长程任务为可执行子任务
- 生成体素轨迹条件,使扩散策略成功率提升44%以上
- 特别适合高噪声或环境变化下的复杂长程操作
扩散策略在机器人操作领域表现优异,但主要局限于短程任务,且在图像噪声下性能显著下降。为此,我们提出一种由视觉语言模型引导的轨迹条件扩散策略(VLM-TDP),实现鲁棒性强、长程的机器人操作。该方法利用先进的视觉语言模型(VLM)将长程任务分解为简洁、可管理的子任务,并创新性地为每个子任务生成体素化轨迹。这些轨迹作为关键条件输入,有效引导扩散策略,显著提升其性能。所提出的轨迹条件扩散策略(TDP)基于示范数据训练,并使用VLM生成的轨迹进行验证。仿真结果表明,该方法显著优于经典扩散策略:平均成功率提升44%,长程任务性能提升超100%,在噪声图像或环境变化等挑战条件下性能下降减少20%。真实世界实验进一步证实,长程任务中性能差距更为明显。视频展示见https://youtu.be/g0T6h32OSC8
原文摘要 · Abstract (English)
Diffusion policy has demonstrated promising performance in the field of robotic manipulation. However, its effectiveness has been primarily limited in short-horizon tasks, and its performance significantly degrades in the presence of image noise. To address these limitations, we propose a VLM-guided trajectory-conditioned diffusion policy (VLM-TDP) for robust and long-horizon manipulation. Specifically, the proposed method leverages state-of-the-art vision-language models (VLMs) to decompose long-horizon tasks into concise, manageable sub-tasks, while also innovatively generating voxel-based trajectories for each sub-task. The generated trajectories serve as a crucial conditioning factor, effectively steering the diffusion policy and substantially enhancing its performance. The proposed Trajectory-conditioned Diffusion Policy (TDP) is trained on trajectories derived from demonstration data and validated using the trajectories generated by the VLM. Simulation experimental results indicate that our method significantly outperforms classical diffusion policies, achieving an average 44% increase in success rate, over 100% improvement in long-horizon tasks, and a 20% reduction in performance degradation in challenging conditions, such as noisy images or altered environments. These findings are further reinforced by our real-world experiments, where the performance gap becomes even more pronounced in long-horizon tasks. Videos are available on https://youtu.be/g0T6h32OSC8
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。