融合可见光与热成像,提升复杂光照下立体深度估计精度
Adaptive Stereo Depth Estimation with Multi-Spectral Images Across All Lighting Conditions
- 将可见光与热成像视为立体图像对,用跨模态匹配构建代价体
- 在光照不良区域用热成像单目深度补全,提升匹配鲁棒性
- 在MS2数据集上达到当前最优,适合低光照场景深度感知应用
恶劣光照下的深度估计仍是重大挑战。近年来,融合可见光与热成像的多光谱深度估计展现出潜力,但现有方法在像素级特征匹配上存在困难,难以充分利用不同波段间的几何约束。为此,我们提出一种新框架,引入立体深度估计以强化几何一致性。具体地,将可见光与热成像视为一对立体图像,通过跨模态特征匹配(CFM)模块构建像素级匹配的代价体。为缓解光照劣化对立体匹配的影响,提出退化掩码机制,利用鲁棒的单目热成像深度估计在退化区域进行补充。该方法在多光谱立体(MS2)数据集上达到当前最优(SOTA)性能,定性评估显示其在多种光照条件下均能生成高质量深度图。
原文摘要 · Abstract (English)
Depth estimation under adverse conditions remains a significant challenge. Recently, multi-spectral depth estimation, which integrates both visible light and thermal images, has shown promise in addressing this issue. However, existing algorithms struggle with precise pixel-level feature matching, limiting their ability to fully exploit geometric constraints across different spectra. To address this, we propose a novel framework incorporating stereo depth estimation to enforce accurate geometric constraints. In particular, we treat the visible light and thermal images as a stereo pair and utilize a Cross-modal Feature Matching (CFM) Module to construct a cost volume for pixel-level matching. To mitigate the effects of poor lighting on stereo matching, we introduce Degradation Masking, which leverages robust monocular thermal depth estimation in degraded regions. Our method achieves state-of-the-art (SOTA) performance on the Multi-Spectral Stereo (MS2) dataset, with qualitative evaluations demonstrating high-quality depth maps under varying lighting conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。