无需配准图像也能精准检测红外可见光目标,适用于真实复杂场景。
Breaking Alignment Barriers: TPS-Driven Semantic Correlation Learning for Alignment-Free RGB-T Salient Object Detection
- 用薄板样条和语义关联模块解决跨模态空间错位问题。
- 在未对齐数据集上达到当前轻量级方法最优性能,比主流方法更鲁棒。
- 适合需要低算力部署的实时多模态目标检测应用。
现有RGB-T显著目标检测方法依赖人工对齐标注数据,在真实世界中面对未经对齐的红外与可见光图像对时表现急剧下降。由于存在空间错位、尺度差异和视角变化等跨模态差异,现有方法在未对齐数据集上性能严重受限。为此,本文提出一种面向真实未对齐图像对的高效RGB-T SOD方法——基于薄板样条的语义关联学习网络(TPS-SCL)。采用双流MobileViT作为编码器,结合高效的Mamba扫描机制,以低参数量和计算开销建模跨模态关联。设计语义关联约束模块(SCCM)分层抑制冗余背景干扰;引入薄板样条对齐模块(TPSAM)缓解模态间空间偏差;并加入跨模态关联模块(CMCM)充分挖掘与融合模态间依赖关系。大量实验表明,TPS-SCL在各类数据集上均达到现有轻量级方法的最先进水平,且优于主流RGB-T SOD方法。
原文摘要 · Abstract (English)
Existing RGB-T salient object detection methods predominantly rely on manually aligned and annotated datasets, struggling to handle real-world scenarios with raw, unaligned RGB-T image pairs. In practical applications, due to significant cross-modal disparities such as spatial misalignment, scale variations, and viewpoint shifts, the performance of current methods drastically deteriorates on unaligned datasets. To address this issue, we propose an efficient RGB-T SOD method for real-world unaligned image pairs, termed Thin-Plate Spline-driven Semantic Correlation Learning Network (TPS-SCL). We employ a dual-stream MobileViT as the encoder, combined with efficient Mamba scanning mechanisms, to effectively model correlations between the two modalities while maintaining low parameter counts and computational overhead. To suppress interference from redundant background information during alignment, we design a Semantic Correlation Constraint Module (SCCM) to hierarchically constrain salient features. Furthermore, we introduce a Thin-Plate Spline Alignment Module (TPSAM) to mitigate spatial discrepancies between modalities. Additionally, a Cross-Modal Correlation Module (CMCM) is incorporated to fully explore and integrate inter-modal dependencies, enhancing detection performance. Extensive experiments on various datasets demonstrate that TPS-SCL attains state-of-the-art (SOTA) performance among existing lightweight SOD methods and outperforms mainstream RGB-T SOD approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。