arXiv:2605.16127cs.CV2026-05

用语言模型动态调整相机与激光雷达融合,提升恶劣天气下3D占位预测精度。

WeatherOcc3D: VLM-Assisted Adverse Weather Aware 3D Semantic Occupancy Prediction

论文配图:WeatherOcc3D: VLM-Assisted Adverse Weather Aware 3D Semantic Occupancy Prediction
图 1 · 摘自论文原文
  • 通过文本语义引导传感器融合,自适应调节相机与激光雷达权重。
  • 在nuScenes数据集上,mIoU分别达26.3和21.1,显著优于基线。
  • 适合自动驾驶中复杂天气场景的感知系统设计者使用。

多模态3D语义占位预测通常通过融合摄像头与激光雷达输入来增强鲁棒性,但其性能受环境变化制约。摄像头在低光照下严重退化,激光雷达在强降水时产生显著后向散射噪声。这些恶劣条件引发模态信任问题,静态融合策略无法在特定传感器失效时动态调整权重。为此,我们提出一种基于视觉语言模型(VLM)的框架,利用预训练的CLIP潜在空间,通过语言环境线索指导多传感器融合。采用参数高效适配器将天气相关的文本嵌入与传感器特征对齐,并设计门控机制,将环境不确定性分解为能见度与光照两个因素。该方法使模型可动态调节融合比例——晴天白天优先使用语义摄像头特征,雨夜则转向几何激光雷达先验。在nuScenes数据集上的评估表明,该框架在OccMamba和M-CONet架构上分别实现26.3和21.1的mIoU,显著优于传统基线。

原文摘要 · Abstract (English)

While multi-modal 3D semantic occupancy prediction typically enhances robustness by fusing camera and LiDAR inputs, its effectiveness is fundamentally constrained by environmental variability. Specifically, camera sensors suffer from severe low-light degradation, while LiDAR sensors encounter significant backscatter noise during heavy precipitation. These adverse conditions create a modality trust problem, as static fusion strategies fail to adaptively re-weight inputs when a specific sensor becomes unreliable. To address this, we propose a VLM-assisted framework leveraging the pre-trained CLIP latent space to guide multi-sensor integration via linguistic environmental cues. We utilize a parameter-efficient adapter to align weather-specific text embeddings with sensor features, coupled with a gating strategy that decomposes environmental uncertainty into two factors: visibility and illumination. This enables the model to dynamically modulate the fusion ratio - prioritizing semantic camera features in clear daylight and shifting to geometric LiDAR priors during rainy nights. Evaluations on the nuScenes dataset demonstrate the versatility of our approach, as implementing our proposed framework on the OccMamba and M-CONet architectures achieves mIoU scores of 26.3 and 21.1, respectively, significantly outperforming their traditional baselines.

3D占位多模态融合恶劣天气视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。