arXiv:2510.03689cs.CV2025-10中稿 · TMM被引 6

解决红外与可见光图像融合中的模态不平衡问题,提升显著目标检测精度。

SAMSOD: Rethinking SAM Optimization for RGB-T Salient Object Detection

论文配图:SAMSOD: Rethinking SAM Optimization for RGB-T Salient Object Detection
图 1 · 摘自论文原文
  • 用单模态监督增强非主导模态学习,缓解双模态训练失衡
  • 通过梯度解冲突机制减少高低激活神经元的梯度干扰
  • 适合多模态图像分割、工业缺陷检测等场景

RGB-T 显著目标检测(SOD)旨在结合可见光与热成像数据分割显著目标。尽管已有研究对 Segment Anything Model(SAM)进行微调以提升性能,但双模态收敛不平衡及高-低激活神经元间显著的梯度差异被忽视,限制了进一步优化空间。本文提出 SAMSOD 模型,利用单模态监督增强非主导模态的学习能力,并引入梯度解冲突策略,降低冲突梯度对模型收敛的影响。同时,采用两个解耦适配器分别掩码高-低激活神经元,通过强化背景学习来突出前景目标。在 RGB-T SOD 基准数据集上的基础实验,以及在草图监督的 RGB-T SOD、全监督的 RGB-D SOD 数据集和全监督的 RGB-D 铁轨表面缺陷检测任务上的泛化性实验,均验证了该方法的有效性。

原文摘要 · Abstract (English)

RGB-T salient object detection (SOD) aims to segment attractive objects by combining RGB and thermal infrared images. To enhance performance, the Segment Anything Model has been fine-tuned for this task. However, the imbalance convergence of two modalities and significant gradient difference between high- and low- activations are ignored, thereby leaving room for further performance enhancement. In this paper, we propose a model called \textit{SAMSOD}, which utilizes unimodal supervision to enhance the learning of non-dominant modality and employs gradient deconfliction to reduce the impact of conflicting gradients on model convergence. The method also leverages two decoupled adapters to separately mask high- and low-activation neurons, emphasizing foreground objects by enhancing background learning. Fundamental experiments on RGB-T SOD benchmark datasets and generalizability experiments on scribble supervised RGB-T SOD, fully supervised RGB-D SOD datasets and full-supervised RGB-D rail surface defect detection all demonstrate the effectiveness of our proposed method.

多模态目标检测图像融合SAM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。