用多光谱图像检测隐身物体,效果优于传统方法
Seeing the Unseen: Camouflaged Object Detection Beyond the Visible Spectrum

- 提出MSFormer框架,直接输入多光谱图像生成掩码
- 在COD任务上实现更优性能,显著提升检测精度
- 适合需要突破可见光限制的低可见场景应用
近年来,隐身物体检测(COD)在低可见度场景中取得显著进展,早期研究已成功定位伪装场景中的目标。然而,现有方法主要依赖传统的三通道RGB图像,导致可用视觉信息局限于有限光谱范围。多光谱图像通过捕捉精细的光谱特征,提供更丰富的场景信息。为此,我们提出一种新方法,利用多光谱图像进行COD,构建了一个端到端框架——MSFormer,以多光谱伪装图像为输入,输出二值掩码。此外,我们还提供了多光谱波段融合在该复杂低视觉任务中的实证依据。大量实验表明,该方法显著优于现有技术。
原文摘要 · Abstract (English)
Recent advances in camouflaged object detection (COD) have led to substantial progress in challenging low-visibility scenarios, with pioneering studies demonstrating notable success in localizing objects in camouflaged scenes. Despite these achievements, existing approaches predominantly rely on conventional three-channel RGB imagery, thereby constraining the available visual information to a limited spectral range. Multispectral images offer a wide range of information about a scene by capturing fine-grained spectral signatures. Hence, by leveraging multispectral images for COD, we introduce a novel approach to detect camouflaged objects from the corresponding multispectral inputs. In particular, we propose an end-to-end framework, \textbf{\textit{MSFormer}}, that takes a multispectral camouflaged image as input and predicts a binary mask for it. Additionally, we also provide empirical justification for integrating multispectral bands for this complex low-vision task. Our extensive experiments demonstrate the effectiveness of our method, which outperforms existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。