提出三流网络提升红外与可见光图像显著目标检测的鲁棒性。
Divide-and-Conquer: Confluent Triple-Flow Network for RGB-T Salient Object Detection
- 采用分治策略,分设可见光、红外和融合三路并行处理
- 在多个数据集上超越现有方法,尤其在噪声干扰下表现更稳
- 适合需要高鲁棒性的跨模态视觉任务开发者参考
RGB-热成像显著目标检测旨在识别对齐的可见光与热红外图像中的显著目标。传统编码器-解码器结构虽设计用于跨模态特征交互,但对缺陷模态带来的噪声鲁棒性不足。受人类视觉系统层级结构启发,本文提出ConTriNet——一种采用分治策略的鲁棒融合三流网络。ConTriNet包含两条模态专用流(分别处理RGB与热成像)和一条模态互补流,融合双模态信息。其核心优势包括:在共享编码器中引入模态诱导特征调制模块,减少模态间差异,降低缺陷样本影响;在分离流中使用基础残差空洞空间金字塔模块,扩大感受野以捕捉多尺度上下文信息;在互补流中设计模态感知动态聚合模块,自适应融合双流显著线索。通过流协同融合策略,进一步优化各流输出的显著图,生成高质量全分辨率预测结果。为评估方法鲁棒性,我们构建了涵盖多种真实挑战场景的综合基准VT-IMAG。在公开数据集及VT-IMAG上的大量实验表明,ConTriNet在常规与挑战性场景中均持续优于当前最优方法。
原文摘要 · Abstract (English)
RGB-Thermal Salient Object Detection aims to pinpoint prominent objects within aligned pairs of visible and thermal infrared images. Traditional encoder-decoder architectures, while designed for cross-modality feature interactions, may not have adequately considered the robustness against noise originating from defective modalities. Inspired by hierarchical human visual systems, we propose the ConTriNet, a robust Confluent Triple-Flow Network employing a Divide-and-Conquer strategy. Specifically, ConTriNet comprises three flows: two modality-specific flows explore cues from RGB and Thermal modalities, and a third modality-complementary flow integrates cues from both modalities. ConTriNet presents several notable advantages. It incorporates a Modality-induced Feature Modulator in the modality-shared union encoder to minimize inter-modality discrepancies and mitigate the impact of defective samples. Additionally, a foundational Residual Atrous Spatial Pyramid Module in the separated flows enlarges the receptive field, allowing for the capture of multi-scale contextual information. Furthermore, a Modality-aware Dynamic Aggregation Module in the modality-complementary flow dynamically aggregates saliency-related cues from both modality-specific flows. Leveraging the proposed parallel triple-flow framework, we further refine saliency maps derived from different flows through a flow-cooperative fusion strategy, yielding a high-quality, full-resolution saliency map for the final prediction. To evaluate the robustness and stability of our approach, we collect a comprehensive RGB-T SOD benchmark, VT-IMAG, covering various real-world challenging scenarios. Extensive experiments on public benchmarks and our VT-IMAG dataset demonstrate that ConTriNet consistently outperforms state-of-the-art competitors in both common and challenging scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。