arXiv:2505.15491cs.CV2025-05中稿 · ICIP 2025被引 5

通过频谱分析融合可见光与热成像,提升低光环境下的语义分割精度。

Spectral-Aware Global Fusion for RGB-Thermal Semantic Segmentation

  • 从频谱角度区分高低频特征,分别处理跨模态差异。
  • 在MFNet和PST900数据集上优于现有方法,显著提升分割鲁棒性。
  • 适合自动驾驶等需多模态感知的高可靠性场景。

仅依赖可见光(RGB)数据的语义分割在低光照、视线遮挡等复杂条件下表现不佳,限制了其在自动驾驶等关键应用中的可靠性。引入热辐射数据与RGB图像融合可增强性能与鲁棒性,但如何有效对齐模态差异并融合特征仍是难题。本文从新的频谱视角出发,观察到多模态特征可分为两类:低频特征提供整体场景上下文(如颜色变化、平滑区域),高频特征则捕捉模态特异性细节(如边缘、纹理)。受此启发,提出谱感知全局融合网络(SGFNet),显式建模高频特征间的跨模态交互,以增强和融合多模态特征。实验表明,SGFNet在MFNet和PST900数据集上均优于当前最优方法。

原文摘要 · Abstract (English)

Semantic segmentation relying solely on RGB data often struggles in challenging conditions such as low illumination and obscured views, limiting its reliability in critical applications like autonomous driving. To address this, integrating additional thermal radiation data with RGB images demonstrates enhanced performance and robustness. However, how to effectively reconcile the modality discrepancies and fuse the RGB and thermal features remains a well-known challenge. In this work, we address this challenge from a novel spectral perspective. We observe that the multi-modal features can be categorized into two spectral components: low-frequency features that provide broad scene context, including color variations and smooth areas, and high-frequency features that capture modality-specific details such as edges and textures. Inspired by this, we propose the Spectral-aware Global Fusion Network (SGFNet) to effectively enhance and fuse the multi-modal features by explicitly modeling the interactions between the high-frequency, modality-specific features. Our experimental results demonstrate that SGFNet outperforms the state-of-the-art methods on the MFNet and PST900 datasets.

多模态融合语义分割热成像频谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。