arXiv:2602.07343cs.CVcs.AI2026-02

根据光照条件动态融合可见光与热成像,提升恶劣环境下的道路分割精度。

Seeing Roads Through Words: A Language-Guided Framework for RGB-T Driving Scene Segmentation

  • 用视觉语言模型指导动态融合策略,按光照状态调节双模态贡献。
  • 在MFNet数据集上达到62.3% mIoU和77.5% mAcc,刷新当前最优性能。
  • 适合需要夜间或阴影环境下高精度语义分割的自动驾驶系统使用。

在恶劣光照、照明和阴影条件下实现鲁棒的道路场景语义分割,仍是自动驾驶的核心挑战。可见光-热成像融合是常用方法,但现有方法对所有条件采用静态融合策略,导致模态特异性噪声在网络中传播。为此,我们提出CLARITY,该方法基于视觉-语言模型(VLM)先验,动态适应检测到的场景光照状态,调节各模态贡献,而非采用固定融合策略。同时引入两种机制:一是保留以往噪声抑制方法误删的有效暗物体语义;二是采用分层解码器,在多尺度间强制结构一致性,以增强细小物体边界清晰度。在MFNet数据集上的实验表明,CLARITY达到了新的最佳性能,取得62.3% mIoU和77.5% mAcc。

原文摘要 · Abstract (English)

Robust semantic segmentation of road scenes under adverse illumination, lighting, and shadow conditions remain a core challenge for autonomous driving applications. RGB-Thermal fusion is a standard approach, yet existing methods apply static fusion strategies uniformly across all conditions, allowing modality-specific noise to propagate throughout the network. Hence, we propose CLARITY that dynamically adapts its fusion strategy to the detected scene condition. Guided by vision-language model (VLM) priors, the network learns to modulate each modality's contribution based on the illumination state while leveraging object embeddings for segmentation, rather than applying a fixed fusion policy. We further introduce two mechanisms - one which preserves valid dark-object semantics that prior noise-suppression methods incorrectly discard, and a hierarchical decoder that enforces structural consistency across scales to sharpen boundaries on thin objects. Experiments on the MFNet dataset demonstrate that CLARITY establishes a new state-of-the-art (SOTA), achieving 62.3% mIoU and 77.5% mAcc.

语义分割多模态融合自动驾驶热成像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。