用热成像引导视觉模型,提升复杂场景下目标检测的鲁棒性。
KAN-SAM: Kolmogorov-Arnold Network Guided Segment Anything Model for RGB-T Salient Object Detection
- 通过KAN网络将热成像作为提示,增强RGB图像表征
- 在多个基准上超越当前最优方法,显著提升检测精度
- 适合需要跨模态融合的红外可见光目标检测任务
现有的RGB-热成像显著目标检测(RGB-T SOD)方法旨在通过融合可见光与热成像信息,在复杂场景中实现稳健的目标识别,但常受限于数据集多样性不足及多模态表示构建效率低下,导致泛化能力有限。本文提出一种基于提示学习的新型RGB-T SOD方法——KAN-SAM,充分利用视觉基础模型的潜力。具体而言,我们通过高效的柯尔莫戈洛夫-阿诺德网络(KAN)适配器,将热成像特征作为引导提示注入Segment Anything Model 2(SAM2),有效增强可见光表征并提升鲁棒性。此外,引入互斥随机掩码策略,降低对可见光数据的依赖,进一步改善泛化性能。在多个基准测试上的实验结果表明,该方法优于现有最先进方法。
原文摘要 · Abstract (English)
Existing RGB-thermal salient object detection (RGB-T SOD) methods aim to identify visually significant objects by leveraging both RGB and thermal modalities to enable robust performance in complex scenarios, but they often suffer from limited generalization due to the constrained diversity of available datasets and the inefficiencies in constructing multi-modal representations. In this paper, we propose a novel prompt learning-based RGB-T SOD method, named KAN-SAM, which reveals the potential of visual foundational models for RGB-T SOD tasks. Specifically, we extend Segment Anything Model 2 (SAM2) for RGB-T SOD by introducing thermal features as guiding prompts through efficient and accurate Kolmogorov-Arnold Network (KAN) adapters, which effectively enhance RGB representations and improve robustness. Furthermore, we introduce a mutually exclusive random masking strategy to reduce reliance on RGB data and improve generalization. Experimental results on benchmarks demonstrate superior performance over the state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。