让视觉语言动作模型具备3D空间与实例理解能力,提升机器人操作精度。
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding

- 通过多视角一致性约束和几何聚合,强化3D特征编码
- 引入高阶实例令牌,实现3D实例感知与遮挡下推理
- 无需额外标注,可直接接入现有模型,适合机器人控制研究
视觉-语言-动作模型在机器人操作中取得显著进展,但缺乏对3D场景的深入理解,主要表现为三方面挑战:3D空间位置提取弱且未强制多视角一致性、3D实例理解不足、遮挡下推理脆弱。尽管已有成熟3D感知方法,但其直接集成到VLA流程中受限于架构不兼容及对昂贵实例级标注的依赖。为此,我们提出3DVLA,一种即插即用框架,在不需额外人工标注且不丢失视觉语言模型先验的前提下,注入鲁棒的3D推理能力。具体包括:(1) 全局3D特征编码结合跨模态多视角一致性约束与空间条件几何聚合;(2) 基于高层实例令牌的实例估计模块,增强3D实例感知;(3) 遮蔽自监督3D编码分支,保留视觉补全预测器以应对遮挡。我们将3DVLA集成至多个VLA基线,在LIBERO-Plus和RoboTwin 2.0上评估,结果表明操纵性能持续且显著提升,验证了方法的有效性与即插即用兼容性。
原文摘要 · Abstract (English)
Vision-Language-Action models have achieved remarkable progress in robotic manipulation, yet they suffer from a critical limitation: a lack of 3D scene understanding. This deficiency manifests as three intertwined challenges: weak extraction of 3D spatial positions without enforcing multi-view consistency, inadequate 3D instance understanding, and fragile reasoning under occlusion. Although mature 3D perception methods exist, their direct integration into VLA pipelines is hindered by architectural incompatibility and by heavy reliance on costly instance-level annotations. To address the above challenges, we propose 3DVLA, a plug-and-play framework that injects robust 3D reasoning into pretrained VLAs without requiring extra manual labels or discarding VLM priors. Specifically, 3DVLA tackles the three challenges through: (1) pervasive 3D feature encoding with explicit multi-view consistency constraints across all modalities and a Spatially-Conditioned Geometry Aggregation method, (2) an instance estimation module with high-level instance tokens for 3D instance awareness, and (3) a masked self-supervised 3D encoding branch that retains its predictor for visual token completion to handle occlusions. We integrate 3DVLA with multiple VLA baselines and evaluate on LIBERO-Plus and RoboTwin 2.0. Results show consistent and significant gains in manipulation performance, validating both the effectiveness and plug-and-play compatibility of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。