用2D图像直接推3D空间结构,让模型自动学会精准定位。
OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
- 把3D占用图当隐式监督信号,从2D图像学3D结构
- 在nuScenes上轨迹预测性能达最新水平,3D问答更准
- 推理时可跳过3D计算,速度不降还更省资源
多模态大语言模型虽具备强大视觉语言推理能力,但在自动驾驶中仍缺乏稳健的三维空间理解。这一局限源于两大挑战:(1) 缺乏低成本且有效的三维表示构建方法,尤其缺少昂贵的人工标注;(2) 视觉语言模型因缺乏大规模三维视觉-语言预训练,导致细粒度空间信息丢失。为此,我们提出OccVLA,一种将三维占用表示融入统一多模态推理过程的新框架。与以往依赖显式3D输入的方法不同,OccVLA将密集3D占用图视为预测输出和监督信号,使模型能直接从二维视觉输入中学习细粒度空间结构。该占用预测被视为隐式推理过程,推理时可跳过而不影响性能,无额外计算开销。OccVLA在nuScenes基准上实现轨迹规划的最先进结果,并在3D视觉问答任务中表现优异,提供了一种可扩展、可解释、全视觉驱动的自动驾驶解决方案。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have shown strong vision-language reasoning abilities but still lack robust 3D spatial understanding, which is critical for autonomous driving. This limitation stems from two key challenges: (1) the difficulty of constructing accessible yet effective 3D representations without expensive manual annotations, and (2) the loss of fine-grained spatial details in VLMs due to the absence of large-scale 3D vision-language pretraining. To address these challenges, we propose OccVLA, a novel framework that integrates 3D occupancy representations into a unified multimodal reasoning process. Unlike prior approaches that rely on explicit 3D inputs, OccVLA treats dense 3D occupancy as both a predictive output and a supervisory signal, enabling the model to learn fine-grained spatial structures directly from 2D visual inputs. The occupancy predictions are regarded as implicit reasoning processes and can be skipped during inference without performance degradation, thereby adding no extra computational overhead. OccVLA achieves state-of-the-art results on the nuScenes benchmark for trajectory planning and demonstrates superior performance on 3D visual question-answering tasks, offering a scalable, interpretable, and fully vision-based solution for autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。