用稀疏查询融合视觉语言模型与占据表示,实现4D场景理解与规划统一。
SparseOccVLA: Bridging Occupancy and Vision-Language Models via Sparse Queries for Unified 4D Scene Understanding and Planning
- 通过稀疏占据编码生成紧凑查询,连接视觉语言与空间占据
- 在OmniDrive-nuScenes上CIDEr提升7%,Occ3D-nuScenes mIoU提高0.5
- 适合自动驾驶中需联合感知、预测与决策的系统开发
在自动驾驶中,视觉语言模型(VLM)擅长高层推理,而语义占据提供精细的空间细节。尽管各领域进展显著,但尚未有方法能有效整合二者。传统VLM存在令牌爆炸和时空推理受限问题,而语义占据虽具统一显式空间表征,却因过密难以高效融入VLM。为此,我们提出SparseOccVLA,一种新型视觉-语言-动作模型,通过稀疏占据查询统一实现场景理解、占据预测与轨迹规划。该模型采用轻量级稀疏占据编码器生成紧凑且信息丰富的稀疏占据查询,作为视觉与语言之间的单一桥梁。这些查询被对齐至语言空间,并由大语言模型(LLM)进行统一推理,完成场景理解与未来占据预测。此外,我们设计了基于LLM引导的锚点扩散规划器,包含解耦的锚点评分与去噪机制,以及跨模型轨迹条件融合。SparseOccVLA在OmniDrive-nuScenes上相对最优模型的CIDEr提升7%,在Occ3D-nuScenes上mIoU提升0.5,并在nuScenes基准上达到最先进的开环规划指标,展现出强大的整体能力。
原文摘要 · Abstract (English)
In autonomous driving, Vision Language Models (VLMs) excel at high-level reasoning , whereas semantic occupancy provides fine-grained details. Despite significant progress in individual fields, there is still no method that can effectively integrate both paradigms. Conventional VLMs struggle with token explosion and limited spatiotemporal reasoning, while semantic occupancy provides a unified, explicit spatial representation but is too dense to integrate efficiently with VLMs. To address these challenges and bridge the gap between VLMs and occupancy, we propose SparseOccVLA, a novel vision-language-action model that unifies scene understanding, occupancy forecasting, and trajectory planning powered by sparse occupancy queries. Starting with a lightweight Sparse Occupancy Encoder, SparseOccVLA generates compact yet highly informative sparse occupancy queries that serve as the single bridge between vision and language. These queries are aligned into the language space and reasoned by the LLM for unified scene understanding and future occupancy forecasting. Furthermore, we introduce an LLM-guided Anchor-Diffusion Planner featuring decoupled anchor scoring and denoising, as well as cross-model trajectory-condition fusion. SparseOccVLA achieves a 7% relative improvement in CIDEr over the state-of-the-art on OmniDrive-nuScenes, a 0.5 increase in mIoU score on Occ3D-nuScenes, and sets state-of-the-art open-loop planning metric on nuScenes benchmark, demonstrating its strong holistic capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。