用3D场景流做运动先验,提升机器人视觉语言动作模型的泛化能力。
LaMP: Learning Vision-Language-Action Policy with 3D Scene Flow as Latent Motion Prior
- 双专家架构:运动专家生成3D场景流,动作专家据此预测行为
- 在三个仿真和真实场景中表现优于现有模型,尤其在未知动态下提升9.7%
- 适合研究机器人具身智能与多模态动作规划的学者
我们提出LaMP,一种双专家视觉-语言-动作框架,将密集3D场景流作为潜在运动先验用于机器人操作。现有VLA模型直接从2D语义视觉特征回归动作,需隐式学习复杂的3D物理交互,导致在陌生空间动态下性能下降。LaMP通过门控交叉注意力将运动匹配的'运动专家'与策略预测的'动作专家'对齐。具体地,运动专家生成一步部分去噪的3D场景流,其隐藏状态仅条件化动作专家,无需完整多步重建。我们在LIBERO、LIBERO-Plus和SimplerEnv-WidowX仿真基准以及真实世界实验中评估了LaMP。结果表明,它在所有基准上均优于对比基线,且在相同训练预算下达到最高平均成功率。在LIBERO-Plus的分布外扰动测试中,相比最强基线平均提升9.7%。
原文摘要 · Abstract (English)
We introduce \textbf{LaMP}, a dual-expert Vision-Language-Action framework that embeds dense 3D scene flow as a latent motion prior for robotic manipulation.Existing VLA models regress actions directly from 2D semantic visual features, forcing them to learn complex 3D physical interactions implicitly.This implicit learning strategy degrades under unfamiliar spatial dynamics.LaMP addresses this limitation by aligning a flow-matching \emph{Motion Expert} with a policy-predicting \emph{Action Expert} through gated cross-attention.Specifically, the Motion Expert generates a one-step partially denoised 3D scene flow, and its hidden states condition the Action Expert without full multi-step reconstruction.We evaluate LaMP on the LIBERO, LIBERO-Plus, and SimplerEnv-WidowX simulation benchmarks as well as real-world experiments.LaMP consistently outperforms evaluated VLA baselines across LIBERO, LIBERO-Plus, and SimplerEnv-WidowX benchmarks, achieving the highest reported average success rates under the same training budgets. On LIBERO-Plus OOD perturbations, LaMP shows improved robustness with an average 9.7\% gain over the strongest prior baseline.Our project page is available at https://summerwxk.github.io/lamp-project-page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。