用最优传输加速生成机器人动作,1-2步完成且速度提升10倍。
Fast Flow-based Visuomotor Policies via Conditional Optimal Transport Couplings
- 通过条件最优传输构建直线型流方程,跳过复杂积分。
- 相比扩散策略成功率高4%,真实任务中1-2步生成优质轨迹。
- 适合需要快速响应的机器人实时控制场景。
扩散和流匹配策略在机器人应用中表现优异,能准确捕捉多模态轨迹分布。但其推理过程依赖常微分方程(ODE)或随机微分方程(SDE)的数值积分,计算开销大,难以作为实时控制器使用。本文提出一种基于条件最优传输耦合的方法,在噪声与样本之间建立直接映射,使流方程的解为直线路径,从而加速动作生成。发现直接耦合在条件任务中失效,因此将条件变量引入耦合过程,显著提升少步预测性能。所提少步策略在多样仿真任务中成功率达4%更高,且推理速度提升10倍;在真实机器人任务中,仅需1-2步即可生成高质量、多样化的动作轨迹。该方法训练复杂度与扩散策略及基础流匹配一致,无需知识蒸馏。
原文摘要 · Abstract (English)
Diffusion and flow matching policies have recently demonstrated remarkable performance in robotic applications by accurately capturing multimodal robot trajectory distributions. However, their computationally expensive inference, due to the numerical integration of an ODE or SDE, limits their applicability as real-time controllers for robots. We introduce a methodology that utilizes conditional Optimal Transport couplings between noise and samples to enforce straight solutions in the flow ODE for robot action generation tasks. We show that naively coupling noise and samples fails in conditional tasks and propose incorporating condition variables into the coupling process to improve few-step performance. The proposed few-step policy achieves a 4% higher success rate with a 10x speed-up compared to Diffusion Policy on a diverse set of simulation tasks. Moreover, it produces high-quality and diverse action trajectories within 1-2 steps on a set of real-world robot tasks. Our method also retains the same training complexity as Diffusion Policy and vanilla Flow Matching, in contrast to distillation-based approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。