用热成像合成短波红外图像,提升恶劣环境下的多模态融合效果。
SWIR-LightFusion: Multi-spectral Semantic Fusion of Synthetic SWIR with Thermal IR (LWIR/MWIR) and RGB
- 从热成像数据生成合成短波红外图像,无需真实SWIR数据
- 在多个公开与私有数据集上显著提升融合图像的对比度和结构保真度
- 适合做智能监控与自动驾驶的视觉系统研发人员参考
恶劣能见度条件下提升场景理解能力仍是监控与自主导航系统的关键挑战。传统可见光(RGB)与热红外(MWIR/LWIR)图像融合常因大气干扰或光照不足而信息不全。短波红外(SWIR)因其可穿透大气、增强材料区分度而成为有前景的模态,但公开的SWIR数据集稀缺。为此,本研究提出从现有LWIR数据出发,通过先进对比度增强技术合成具有SWIR结构/对比度特征的图像(不追求光谱还原)。进一步构建融合合成SWIR、LWIR与RGB的多模态框架,采用带模态专用编码器与softmax门控融合头的优化编码器-解码器网络。在公开基准(M3FD、TNO、CAMEL、MSRS、RoadScene)及一个私有真实RGB-MWIR-SWIR数据集上的实验表明,该合成SWIR增强融合框架在保持实时性能的同时,显著提升融合图像质量(对比度、边缘清晰度、结构保真度)。还引入公平的三模态基线(LP、LatLRR、GFF)与级联式三模态变体(U2Fusion/SwinFusion),统一评估协议。结果表明该方法在监控与自动驾驶等实际场景中具备显著应用潜力。
原文摘要 · Abstract (English)
Enhancing scene understanding in adverse visibility conditions remains a critical challenge for surveillance and autonomous navigation systems. Conventional imaging modalities, such as RGB and thermal infrared (MWIR / LWIR), when fused, often struggle to deliver comprehensive scene information, particularly under conditions of atmospheric interference or inadequate illumination. To address these limitations, Short-Wave Infrared (SWIR) imaging has emerged as a promising modality due to its ability to penetrate atmospheric disturbances and differentiate materials with improved clarity. However, the advancement and widespread implementation of SWIR-based systems face significant hurdles, primarily due to the scarcity of publicly accessible SWIR datasets. In response to this challenge, our research introduces an approach to synthetically generate SWIR-like structural/contrast cues (without claiming spectral reproduction) images from existing LWIR data using advanced contrast enhancement techniques. We then propose a multimodal fusion framework integrating synthetic SWIR, LWIR, and RGB modalities, employing an optimized encoder-decoder neural network architecture with modality-specific encoders and a softmax-gated fusion head. Comprehensive experiments on public RGB-LWIR benchmarks (M3FD, TNO, CAMEL, MSRS, RoadScene) and an additional private real RGB-MWIR-SWIR dataset demonstrate that our synthetic-SWIR-enhanced fusion framework improves fused-image quality (contrast, edge definition, structural fidelity) while maintaining real-time performance. We also add fair trimodal baselines (LP, LatLRR, GFF) and cascaded trimodal variants of U2Fusion/SwinFusion under a unified protocol. The outcomes highlight substantial potential for real-world applications in surveillance and autonomous systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。