构建2万对无配准红外可见光数据集,提出渐进相关网络实现精准目标检测。
Alignment-Free RGB-T Salient Object Detection: A Large-scale Dataset and Progressive Correlation Network
- 无需图像对齐,通过显式建模跨模态与单模态相关性进行检测
- 在UVT20K数据集上达到新基准,提升复杂场景下检测准确率
- 适合做多模态视觉、红外成像及无人系统感知研究者参考
无配准的可见光-热成像显著目标检测(RGB-T SOD)旨在直接利用未对齐图像对中的互补信息,以应对复杂场景下的鲁棒性挑战。然而,现有基准因人工采集与标注成本高,规模受限,制约了该领域发展。本文构建了一个大规模、高多样性的无配准RGB-T SOD数据集UVT20K,包含20,000对图像、407个场景和1256类显著目标。所有样本均来自真实世界,涵盖低照度、图像杂乱、复杂目标等挑战。每张图像配有完整标注:显著性掩码、草图、边界和挑战属性。同时提出渐进相关网络(PCNet),基于显式对齐建模跨模态与单模态相关性,在未对齐图像对上实现精准预测。大量实验表明该方法在无配准与有配准数据集上均表现优异。代码与数据集已开源。
原文摘要 · Abstract (English)
Alignment-free RGB-Thermal (RGB-T) salient object detection (SOD) aims to achieve robust performance in complex scenes by directly leveraging the complementary information from unaligned visible-thermal image pairs, without requiring manual alignment. However, the labor-intensive process of collecting and annotating image pairs limits the scale of existing benchmarks, hindering the advancement of alignment-free RGB-T SOD. In this paper, we construct a large-scale and high-diversity unaligned RGB-T SOD dataset named UVT20K, comprising 20,000 image pairs, 407 scenes, and 1256 object categories. All samples are collected from real-world scenarios with various challenges, such as low illumination, image clutter, complex salient objects, and so on. To support the exploration for further research, each sample in UVT20K is annotated with a comprehensive set of ground truths, including saliency masks, scribbles, boundaries, and challenge attributes. In addition, we propose a Progressive Correlation Network (PCNet), which models inter- and intra-modal correlations on the basis of explicit alignment to achieve accurate predictions in unaligned image pairs. Extensive experiments conducted on unaligned and aligned datasets demonstrate the effectiveness of our method.Code and dataset are available at https://github.com/Angknpng/PCNet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。