arXiv:2608.09202cs.AI2026-08被引 1

用视觉语言模型提升自动驾驶多传感器融合的可靠性

CRUISE: Vision-Language Model-Guided Uncertainty-Aware Cross-Modal Sensor Fusion for Robust Autonomous Driving

论文配图:CRUISE: Vision-Language Model-Guided Uncertainty-Aware Cross-Modal Sensor Fusion for Robust Autonomous Driving
图 1 · 摘自论文原文
  • 用视觉语言模型生成像素级不确定性估计
  • 动态自适应机制捕捉跨模态依赖关系
  • 适合复杂场景下追求高鲁棒性的自动驾驶研究

现代自动驾驶车辆配备多种传感器(如摄像头、激光雷达、雷达)以实现全面环境感知。然而,在不同驾驶条件下(如能见度低、恶劣天气),各传感器的可靠性差异显著,导致跨模态特征融合难以稳健实现。虽然不确定性量化(UQ)可通过优先处理可靠信号缓解此问题,但现有方法通常仅依赖简单的特征级不确定性估计,在复杂、分布外场景中泛化能力不足。为此,我们提出CRUISE——一种新型不确定性感知的跨模态传感器融合框架。CRUISE引入视觉语言模型(VLM)引导的不确定性量化模块,生成细粒度的像素级不确定性估计。通过利用VLM丰富的先验知识和出色的上下文推理能力,该方法为融合过程提供高度信息性的指导。此外,我们设计了动态自适应机制,显式建模并捕获跨模态依赖关系,确保充分挖掘多传感器输入的内在互补性。

原文摘要 · Abstract (English)

Modern autonomous vehicles are equipped with multiple sensors, such as cameras, LiDAR, and radar, for comprehensive environmental perception. However, robust cross-modal feature fusion remains a critical challenge, as the reliability of each sensor varies significantly across diverse real-world driving conditions, including poor visibility and adverse weather. While uncertainty quantification (UQ) mitigates this issue by allowing models to prioritize reliable signals, existing uncertainty-aware fusion methods typically rely on simple feature-level uncertainty estimates and thus often fail to generalize effectively in complex, out-of-distribution scenarios. To address this limitation, we propose CRUISE, a novel uncertainty-aware cross-modal sensor fusion framework. CRUISE integrates a vision-language model (VLM)-guided UQ module that generates fine-grained, pixel-level uncertainty estimates. By leveraging the VLM's rich prior knowledge and superior contextual reasoning, our approach provides a highly informative guide for the fusion process. Furthermore, we introduce a dynamic adaptive mechanism that explicitly models and captures cross-modal dependencies, ensuring the framework fully exploits the inherent complementary nature of multi-sensor inputs.

自动驾驶多模态融合不确定性量化视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。