用可学习的先验分布替代固定高斯,提升机器人抓取行为克隆效果。
Flowing With Purpose: Latent Action Guided Flow Matching Policies For Robotic Manipulation

- 基于隐动作模型选择专用先验分布,实现动态初始化。
- 真实场景任务成功率提升23.4%,LIBERO-90上提升10.4%。
- 小模型超越大视觉语言模型,适合资源受限场景。
流匹配已成为机器人操作中行为克隆的新标准。然而,最先进的流匹配策略存在系统性结构不匹配:它们依赖全局固定的各向同性源分布,而机器人动作空间具有强烈碎片化和异方差性。这种无差别初始化迫使模型学习高度纠缠的向量场,限制了训练效率和整体性能。为此,我们提出隐动作引导流匹配(LAFM),用可学习的自适应先验分布库替代单一高斯分布。通过隐动作模型将当前观测映射到离散运动基元,选择与动作结构对齐的专用基分布,为去噪过程提供有指导性的初始化。该动态适应性自然捕捉人类示范中的异方差性,使传输轨迹更短、更少纠缠。实验表明,LAFM显著优于标准流匹配方法,在真实机器人部署中任务成功率提升23.4%,在LIBERO-90基准上提升10.4%。此外,其性能超越大规模预训练视觉-语言-动作模型,且使用更小模型架构。
原文摘要 · Abstract (English)
Flow matching has recently become a new standard for behavior cloning in robotic manipulation. However, state-of-the-art flow matching policies suffer from a systematic structural mismatch: they rely on a globally fixed isotropic source distribution despite the strongly fragmented and heteroscedastic structure of robotic action spaces. This agnostic initialization forces the model to learn highly entangled vector fields, bottlenecking training efficiency and limiting overall policy performance. To address this limitation, we introduce Latent Action Guided Flow Matching (LAFM), a novel framework that replaces the monolithic Gaussian with an adaptive library of learned prior distributions. By grounding these distributions using a latent action model, LAFM maps current observations to discrete motion primitives, selecting a specialized base distribution that provides an informed, structurally aligned initialization for the denoising process. This dynamic adaptivity naturally accommodates heteroscedasticity in human demonstrations and makes transport trajectories shorter and less entangled. Empirically, LAFM substantially outperforms standard flow matching formulations, increasing task success rates by 23.4% in real-world robotic deployments and by 10.4% on the LIBERO-90 benchmark. Furthermore, we demonstrate that LAFM achieves state-of-the-art results, surpassing massively pre-trained vision-language-action models while utilizing significantly smaller architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。