arXiv:2505.12715cs.CV2025-05中稿 · KDD

用视觉语言模型动态调节多传感器权重,提升复杂环境下的目标检测鲁棒性。

VLC Fusion: Vision-Language Conditioned Sensor Fusion for Robust Object Detection

  • 通过视觉语言模型感知光照、雨天等环境因素,动态调整图像、激光雷达和红外数据的融合权重。
  • 在真实自动驾驶与军事目标检测数据集上,新方法在已见与未见场景中均显著提升检测准确率。
  • 适合需要应对复杂多变环境的自动驾驶、安防监控等实际应用。

尽管融合多种传感器可提升目标检测性能,但现有融合方法常忽略环境条件和传感器输入的细微差异,难以自适应地调整各模态权重。为此,我们提出视觉语言条件融合(VLC Fusion),利用视觉语言模型(VLM)将融合过程依赖于精细的环境线索。通过捕捉黑暗、降雨、摄像头模糊等高层环境上下文,VLM引导模型根据当前场景动态调节各模态权重,从而增强对环境变化的鲁棒性。我们在包含图像、激光雷达和中波红外模态的真实世界自动驾驶与军事目标检测数据集上评估了该方法。实验表明,VLC Fusion在已见与未见场景中均持续优于传统融合基线,显著提升了检测准确性。

原文摘要 · Abstract (English)

Although fusing multiple sensor modalities can enhance object detection performance, existing fusion approaches often overlook subtle variations in environmental conditions and sensor inputs. As a result, they struggle to adaptively weight each modality under such variations. To address this challenge, we introduce Vision-Language Conditioned Fusion (VLC Fusion), a novel fusion framework that leverages a Vision-Language Model (VLM) to condition the fusion process on nuanced environmental cues. By capturing high-level environmental context such as darkness, rain, and camera blurring, the VLM guides the model to dynamically adjust modality weights based on the current scene, ensuring robustness against environmental shifts. We evaluate VLC Fusion on real-world autonomous driving and military target detection datasets that include image, LIDAR, and mid-wave infrared modalities. Our experiments show that VLC Fusion consistently outperforms conventional fusion baselines, achieving improved detection accuracy in both seen and unseen scenarios.

多模态融合目标检测视觉语言模型自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。