arXiv:2508.17488cs.CV2025-08

通过重要性正则化优化多模态跟踪器,提升模型适应性与稳定性。

Optimizing Multi-Modality Trackers via Significance-Regularized Tuning

  • 引入参数重要性正则化,动态调节预训练模型在多模态场景下的更新
  • 在多个基准上超越当前最优方法,显著提升跨模态迁移能力
  • 适合需要高效适配多模态数据的视觉跟踪研究者

本文针对多模态跟踪器优化难题,提出一种新颖的重要性正则化微调框架,以有效适配预训练的RGB模型。现有微调范式在灵活性与约束性之间摇摆,导致塑性-稳定性权衡不佳。通过系统研究从预训练到多模态场景的过渡过程,我们发现对保持基础模式和管理跨域变化至关重要的参数是问题核心。首先,利用预训练权重的切空间测量并定向先验重要性,以保护泛化能力;随后,在微调阶段刻画转移重要性,强调适应性与稳定性。将两类重要性作为统一正则项融合,显著增强跨模态可迁移性。大量实验表明,该方法在多个多模态跟踪基准上超越现有最先进水平。代码与模型已公开于 https://github.com/zhiwen-xdu/SRTrack。

原文摘要 · Abstract (English)

This paper tackles the critical challenge of optimizing multi-modality trackers by effectively adapting pre-trained models for RGB data. Existing fine-tuning paradigms oscillate between excessive flexibility and over-restriction, both leading to suboptimal plasticity-stability trade-offs. To mitigate this dilemma, we propose a novel significance-regularized fine-tuning framework, which delicately refines the learning process by incorporating intrinsic parameter significance. Through a comprehensive investigation of the transition from pre-trained to multi-modality contexts, we identify that parameters crucial to preserving foundational patterns and managing cross-domain shifts are the primary drivers of this issue. Specifically, we first probe the tangent space of pre-trained weights to measure and orient prior significance, dedicated to preserving generalization. Subsequently, we characterize transfer significance during the fine-tuning phase, emphasizing adaptability and stability. By incorporating these parameter significance terms as unified regularization, our method markedly enhances transferability across modalities. Extensive experiments showcase the superior performance of our method, surpassing current state-of-the-art techniques across various multi-modal tracking benchmarks. The source code and models are publicly available at https://github.com/zhiwen-xdu/SRTrack.

多模态跟踪微调优化参数重要性视觉追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。