arXiv:2602.10698cs.CVcs.AI2026-02被引 4

用深度信息增强视觉语言动作模型的3D理解能力

AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models

  • 引入深度估计模型提取RGB图像中的3D几何线索
  • 在复杂场景中提升感知准确率与动作预测性能
  • 适合需要3D空间理解的机器人控制研究者

视觉-语言-动作(VLA)模型在机器人感知与控制中取得显著进展,但多数方法依赖于基于2D图像训练的视觉语言模型,限制了其在复杂3D环境中的空间理解与动作定位能力。为此,本文提出一种新框架,将深度估计融入VLA模型以丰富3D特征表示。具体地,采用名为VGGT的深度估计基线从标准RGB输入中提取具有几何感知的3D线索,从而在不改变数据分布的前提下,高效利用大规模2D数据集并隐式恢复3D结构信息。为进一步提升深度导出特征的可靠性,引入动作助手模块,通过动作先验约束学习到的3D表示,确保其与下游控制任务的一致性。通过融合增强后的3D特征与传统2D视觉标记,所提方法显著提升了VLA模型的泛化能力与鲁棒性。实验表明,该方法不仅增强了几何模糊场景下的感知能力,还实现了更优的动作预测准确率。本工作展示了深度驱动的数据增强与辅助专家监督在弥合2D观测与3D感知决策间差距方面的潜力。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have recently achieved remarkable progress in robotic perception and control, yet most existing approaches primarily rely on VLM trained using 2D images, which limits their spatial understanding and action grounding in complex 3D environments. To address this limitation, we propose a novel framework that integrates depth estimation into VLA models to enrich 3D feature representations. Specifically, we employ a depth estimation baseline called VGGT to extract geometry-aware 3D cues from standard RGB inputs, enabling efficient utilization of existing large-scale 2D datasets while implicitly recovering 3D structural information. To further enhance the reliability of these depth-derived features, we introduce a new module called action assistant, which constrains the learned 3D representations with action priors and ensures their consistency with downstream control tasks. By fusing the enhanced 3D features with conventional 2D visual tokens, our approach significantly improves the generalization ability and robustness of VLA models. Experimental results demonstrate that the proposed method not only strengthens perception in geometrically ambiguous scenarios but also leads to superior action prediction accuracy. This work highlights the potential of depth-driven data augmentation and auxiliary expert supervision for bridging the gap between 2D observations and 3D-aware decision-making in robotic systems.

机器人控制3D感知多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。