arXiv:2412.16654cs.CV2024-12被引 2

仅用3%可训练参数,实现红外可见光任务的高效互补学习。

IV-tuning: Parameter-Efficient Transfer Learning for Infrared-Visible Tasks

  • 引入轻量级微调方法IV-tuning,仅激活少量参数。
  • 在多个任务上超越现有最优方法,保持高泛化能力。
  • 适合资源受限场景下跨模态视觉任务部署。

现有红外-可见光(IR-VIS)方法沿用预训练视觉模型(PVMs)的通用表征以促进互补学习。然而,我们分析发现,在全量微调范式下,特征空间高度受限且秩低,严重影响泛化性能。一种解决方案是冻结参数,以保留预训练知识并维持特征多样性。为此,我们提出IV-tuning,用于参数高效地利用PVMs完成多种IR-VIS下游任务,包括显著性目标检测、语义分割和目标检测。大量实验表明,IV-tuning优于先前最先进方法,在泛化性和可扩展性方面表现优异。值得注意的是,仅需单个主干网络和3%可训练主干参数,即可有效实现红外与可见光模态的互补学习,相比传统IR-VIS范式具有更优计算效率。

原文摘要 · Abstract (English)

Existing infrared and visible (IR-VIS) methods inherit the general representations of Pre-trained Visual Models (PVMs) to facilitate complementary learning. However, our analysis indicates that under the full fine-tuning paradigm, the feature space becomes highly constrained and low-ranked, which has been proven to seriously impair generalization. One remedy is to freeze the parameters, which preserves pretrained knowledge and helps maintain feature diversity. To this end, we propose IV-tuning, to parameter-efficiently harness PVMs for various IR-VIS downstream tasks, including salient object detection, semantic segmentation, and object detection. Extensive experiments across various settings demonstrate that IV-tuning outperforms previous state-of-the-art methods, and exhibits superior generalization and scalability. Remarkably, with only a single backbone, IV-tuning effectively facilitates the complementary learning of infrared and visible modalities with merely 3% trainable backbone parameters, and achieves superior computational efficiency compared to conventional IR-VIS paradigms.

跨模态轻量化红外可见光

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。