arXiv:2603.02609cs.CVcs.RO2026-03被引 3

用视觉语言模型提升自动驾驶3D语义占位预测的准确性与鲁棒性

VLMFusionOcc3D: VLM Assisted Multi-Modal 3D Semantic Occupancy Prediction

  • 通过视觉语言模型注入高层语义先验,解决稀疏网格下的语义模糊问题
  • 在nuScenes和SemanticKITTI上显著提升恶劣天气下的预测性能
  • 适合需要高可靠性的自动驾驶场景感知系统研发人员

本文提出VLMFusionOcc3D,一种用于自动驾驶中密集3D语义占位预测的鲁棒多模态框架。现有基于体素的占位模型在稀疏几何网格中常面临语义模糊问题,且在恶劣天气下性能下降。为此,我们利用视觉语言模型(VLM)丰富的语言先验,将模糊的体素特征锚定到稳定的语义概念。框架采用双分支特征提取管道,将多视角图像与激光雷达点云投影至统一体素空间。提出实例驱动的VLM注意力(InstVLM),结合门控交叉注意力与LoRA适配的CLIP嵌入,直接向3D体素注入高层语义与地理先验。进一步引入天气感知自适应融合(WeathFusion),基于车辆元数据与天气条件提示动态重加权传感器贡献。为保证结构一致性,采用深度感知几何对齐(DAGA)损失,对齐相机推导的稠密几何与激光雷达返回的稀疏空间精度。在nuScenes和SemanticKITTI数据集上的大量实验表明,所提即插即用模块持续提升先进体素基基线性能。尤其在挑战性天气场景中表现显著提升,提供了一种可扩展、鲁棒的复杂城市导航解决方案。

原文摘要 · Abstract (English)

This paper introduces VLMFusionOcc3D, a robust multimodal framework for dense 3D semantic occupancy prediction in autonomous driving. Current voxel-based occupancy models often struggle with semantic ambiguity in sparse geometric grids and performance degradation under adverse weather conditions. To address these challenges, we leverage the rich linguistic priors of Vision-Language Models (VLMs) to anchor ambiguous voxel features to stable semantic concepts. Our framework initiates with a dual-branch feature extraction pipeline that projects multi-view images and LiDAR point clouds into a unified voxel space. We propose Instance-driven VLM Attention (InstVLM), which utilizes gated cross-attention and LoRA-adapted CLIP embeddings to inject high-level semantic and geographic priors directly into the 3D voxels. Furthermore, we introduce Weather-Aware Adaptive Fusion (WeathFusion), a dynamic gating mechanism that utilizes vehicle metadata and weather-conditioned prompts to re-weight sensor contributions based on real-time environmental reliability. To ensure structural consistency, a Depth-Aware Geometric Alignment (DAGA) loss is employed to align dense camera-derived geometry with sparse, spatially accurate LiDAR returns. Extensive experiments on the nuScenes and SemanticKITTI datasets demonstrate that our plug-and-play modules consistently enhance the performance of state-of-the-art voxel-based baselines. Notably, our approach achieves significant improvements in challenging weather scenarios, offering a scalable and robust solution for complex urban navigation.

3D占位多模态融合视觉语言模型自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。