让视觉语言动作模型学会3D空间推理与动态预测,提升机器人操作精度。
Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation

- 通过点云对齐2D位置嵌入,实现3D几何直接编码
- 自监督重建当前与预测未来点云,捕捉物理动态
- 分层时序动作建模,生成连贯机械臂动作
近期的视觉-语言-动作(VLA)模型在多样化任务中展现出强大泛化能力。然而,真实环境中的有效机器人操作需依赖几何理解与空间推理。现有方法受限于3D数据稀缺及编码过程中的信息丢失,难以同时建模3D结构与动态动作。为此,我们提出Lift3D-VLA,一个统一的VLA框架,赋予模型显式的3D点云推理能力,并支持时序一致的动作生成。基于前期工作Lift3D,我们改进了2D模型提升策略,将3D点与预训练2D位置嵌入几何对齐,实现点云在视觉编码器中的直接处理,减少空间信息损失。在此基础上,提出几何中心掩码自编码(GC-MAE),通过双重目标自监督学习,重建当前点云并预测其未来演化,使2D视觉编码器内化3D结构与物理动态。为充分挖掘3D表示,进一步设计分层时序动作建模,利用LLM多层协同预测动作片段,实现时序一致性。在22个仿真任务和8个真实世界操作任务中,Lift3D-VLA在MetaWorld和RLBench上分别取得10.8%和11.1%更高的平均成功率,优于最强基线4个百分点,且对分布外扰动具有更强泛化能力。
原文摘要 · Abstract (English)
Recently, Vision-Language-Action (VLA) models have demonstrated strong generalization across diverse tasks. However, effective robotic manipulation in physical environments fundamentally requires geometric understanding and spatial reasoning. While some VLA approaches attempt to incorporate 3D information, they are constrained by limited data availability and geometric information loss in current 3D encoding pipelines, and fail to jointly capture 3D geometry and temporally structured actions in dynamic environments. To address these limitations, we introduce Lift3D-VLA, a unified VLA framework that equips models with explicit 3D point cloud reasoning and enables temporally coherent action generation. First, building upon our previous work Lift3D, an enhanced 2D model-lifting strategy is proposed to geometrically align 3D points with pretrained 2D positional embeddings. This design enables direct point-cloud encoding within the VLA vision encoder while minimizing spatial information loss. Based on explicit 3D inputs, we propose Geometry-Centric Masked Autoencoding (GC-MAE), a dual-objective self-supervised framework that reconstructs the current point cloud while predicting its future geometric evolution. This formulation allows the 2D vision encoder to internalize both 3D structure and physical dynamics. To fully exploit 3D representations, we further design layer-wise temporal action modeling, which leverages multiple layers of the LLM to collaboratively predict action chunks, enabling temporally consistent predictions. Across 22 simulated tasks and 8 real-world manipulation tasks, Lift3D-VLA achieves 10.8% and 11.1% higher mean success rates on MetaWorld and RLBench than the best-performing prior VLA methods, and outperforms the strongest real-world baseline by 4 percentage points, while exhibiting stronger generalization to out-of-distribution perturbations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。