arXiv:2605.11832cs.RO2026-05被引 1

用多视角隐空间先验提升机器人操作的感知与动作预测能力

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation

论文配图:Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation
图 1 · 摘自论文原文
  • 通过多视角扩散模型生成新视角隐表示,结合几何引导门控变压器对齐特征
  • 在LIBERO、RoboTwin 2.0等数据集上实现超过主流方法的完成率与鲁棒性
  • 适合关注视觉-语言-动作模型中动作学习效率与3D几何理解的研究者

本文针对视觉-语言-动作(VLA)模型中的空间感知与操作挑战提出解决方案。为缓解单目输入带来的深度模糊问题,我们利用预训练的多视角扩散模型合成潜在新视图,并提出几何引导门控变换器(G3T),在三维几何引导下对齐多视角特征,同时自适应过滤遮挡噪声。为提升动作学习效率,引入动作流形学习(AML),直接在有效动作流形上预测动作,避免对噪声或速度等无结构目标进行低效回归。在LIBERO、RoboTwin 2.0及真实机器人任务上的实验表明,该方法在成功率与鲁棒性方面均优于当前最优基线。项目页面:https://junjxiao.github.io/Multi-view-VLA.github.io/

原文摘要 · Abstract (English)

This paper tackles spatial perception and manipulation challenges in Vision-Language-Action (VLA) models. To address depth ambiguity from monocular input, we leverage a pre-trained multi-view diffusion model to synthesize latent novel views and propose a Geometry-Guided Gated Transformer (G3T) that aligns multi-view features under 3D geometric guidance while adaptively filtering occlusion noise. To improve action learning efficiency, we introduce Action Manifold Learning (AML), which directly predicts actions on the valid action manifold, bypassing inefficient regression of unstructured targets like noise or velocity. Experiments on LIBERO, RoboTwin 2.0, and real-robot tasks show our method achieves superior success rate and robustness over SOTA baselines. Project page: https://junjxiao.github.io/Multi-view-VLA.github.io/.

机器人操作多视角动作学习扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。