提升机器人视觉语言模型的3D感知能力,通过自适应融合2D语义与3D几何信息。
3D-Mix for VLA: A Plug-and-Play Module for Integrating VGGT-based 3D Information into Vision-Language-Action Models
- 采用语义条件门控融合,动态平衡2D与3D特征。
- 在6个大模型系列上平均提升7.0%的跨域任务表现。
- 可即插即用,适配多种主流视觉语言动作架构。
视觉-语言-动作(VLA)模型利用多模态大语言模型(MLLM)进行机器人控制,但现有研究发现,由于主要在二维数据上训练,MLLM的空间智能有限,导致操作任务中3D感知不足。尽管近期方法引入如VGGT等专用3D视觉模型以增强空间理解,但其集成方式多样且缺乏系统性分析,最优融合策略尚不明确。我们开展了一项全面的初步研究,对比九种VGGT集成方案在标准基准上的表现,发现基于任务上下文自适应调节2D语义与3D几何特征的语义条件门控融合,在所有九种方案中表现最佳。我们提出3D-Mix,一个无需修改原有MLLM或动作专家组件的即插即用模块,可集成至GR00T型与π型等多种VLA架构。在SIMPLER和LIBERO基准上对六种MLLM系列(九个模型变体,参数量2B–8B)的实验表明,3D-Mix在所有九个GR00T型变体上于跨域(OOD)SIMPLER基准平均提升7.0%,为增强VLA系统的空间智能提供了系统化方法。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models leverage Multimodal Large Language Models (MLLMs) for robotic control, but recent studies reveal that MLLMs exhibit limited spatial intelligence due to training predominantly on 2D data, resulting in inadequate 3D perception for manipulation tasks. While recent approaches incorporate specialized 3D vision models such as VGGT to enhance spatial understanding, they employ diverse integration mechanisms without systematic investigation, leaving the optimal fusion strategy unclear. We conduct a comprehensive pilot study comparing nine VGGT integration schemes on standardized benchmarks and find that semantic-conditioned gated fusion, which adaptively balances 2D semantic and 3D geometric features based on task context, achieved the strongest performance among all nine evaluated fusion schemes in our pilot study. We present 3D-Mix, a plug-and-play module that integrates into diverse VLA architectures (GR00T-style and $π$-style) without modifying existing MLLM or action expert components. Experiments across six MLLM series (nine model variants, 2B--8B parameters) on SIMPLER and LIBERO show that 3D-Mix delivers consistent performance gains, averaging +7.0% on the out-of-domain (OOD) SIMPLER benchmark across all nine GR00T-style variants, establishing a principled approach for enhancing spatial intelligence in VLA systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。