arXiv:2503.04821eess.IVcs.AI2025-03被引 2

融合可见光与热红外图像,提升复杂环境下的深度估计精度

RGB-Thermal Infrared Fusion for Robust Depth Estimation in Complex Environments

  • 用双模态融合机制对齐可见光与热红外特征,增强跨模态理解
  • 在夜间、雨天等复杂场景下,深度图质量显著优于单模态方法
  • 适合自动驾驶、机器人导航等需高鲁棒性深度感知的领域

复杂现实场景中的深度估计极具挑战性,尤其当仅依赖可见光或热红外(THR)单一模态时。本文提出一种新型多模态深度估计模型RTFusion,通过融合RGB与THR数据的互补优势,提升深度估计的准确性和鲁棒性。RGB提供丰富的纹理和色彩信息,而THR在极端光照条件下仍能保持稳定性。模型引入独特的EGFusion融合机制,包含用于跨模态特征对齐的互惠互补注意力(MCA)模块,以及增强边缘显著性的边缘细节保留模块(ESEM)。在MS2和ViViD++数据集上的大量实验表明,该方法在夜间、雨天及强眩光等复杂环境下均能持续生成高质量深度图,展现出在自动驾驶、机器人及增强现实等场景中可靠深度感知的潜力。

原文摘要 · Abstract (English)

Depth estimation in complex real-world scenarios is a challenging task, especially when relying solely on a single modality such as visible light or thermal infrared (THR) imagery. This paper proposes a novel multimodal depth estimation model, RTFusion, which enhances depth estimation accuracy and robustness by integrating the complementary strengths of RGB and THR data. The RGB modality provides rich texture and color information, while the THR modality captures thermal patterns, ensuring stability under adverse lighting conditions such as extreme illumination. The model incorporates a unique fusion mechanism, EGFusion, consisting of the Mutual Complementary Attention (MCA) module for cross-modal feature alignment and the Edge Saliency Enhancement Module (ESEM) to improve edge detail preservation. Comprehensive experiments on the MS2 and ViViD++ datasets demonstrate that the proposed model consistently produces high-quality depth maps across various challenging environments, including nighttime, rainy, and high-glare conditions. The experimental results highlight the potential of the proposed method in applications requiring reliable depth estimation, such as autonomous driving, robotics, and augmented reality.

深度估计多模态融合热红外

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。