提出VFP方法,让机器人在复杂操作中更灵活地应对多种行为模式。
VFP: Variational Flow-Matching Policy for Multi-Modal Robot Manipulation
- 用变分潜变量先验引导多模态动作生成,避免行为平均化
- 仿真任务中成功率比基线提升49%,真实机器人任务也表现更优
- 适合需要快速推理和多样化决策的机器人控制场景
基于流匹配的策略近期成为学习型机器人操作的有前途方法,相比扩散模型可显著加速动作采样。然而,传统流匹配方法在处理多模态问题时表现不佳,常在复杂操作任务中退化为平均或模糊的行为。为此,我们提出变分流匹配策略(VFP),引入变分潜变量先验以实现模式感知的动作生成,并有效捕捉任务级与轨迹级的多模态特性。VFP进一步结合柯尔莫哥洛夫最优传输(K-OT)实现分布对齐,并采用专家混合(MoE)解码器实现模式专一性与高效推理。我们在41个仿真任务和3个真实机器人任务上进行全面评估,结果表明VFP在仿真环境中相较标准流基基线任务成功率提升49%,在真实机器人任务中表现更优,同时保持快速推理与紧凑模型规模。
原文摘要 · Abstract (English)
Flow-matching-based policies have recently emerged as a promising approach for learning-based robot manipulation, offering significant acceleration in action sampling compared to diffusion-based policies. However, conventional flow-matching methods struggle with multi-modality, often collapsing to averaged or ambiguous behaviors in complex manipulation tasks. To address this, we propose the Variational Flow-Matching Policy (VFP), which introduces a variational latent prior for mode-aware action generation and effectively captures both task-level and trajectory-level multi-modality. VFP further incorporates Kantorovich Optimal Transport (K-OT) for distribution-level alignment and utilizes a Mixture-of-Experts (MoE) decoder for mode specialization and efficient inference. We comprehensively evaluate VFP on 41 simulated tasks and 3 real-robot tasks, demonstrating its effectiveness and sampling efficiency in both simulated and real-world settings. Results show that VFP achieves a 49% relative improvement in task success rate over standard flow-based baselines in simulation, and further outperforms them on real-robot tasks, while still maintaining fast inference and a compact model size. More details are available on our project page: https://sites.google.com/view/varfp/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。