通过频域提示增强,提升红外与可见光追踪的鲁棒性
Robust RGB-T Tracking via Learnable Visual Fourier Prompt Fine-tuning and Modality Fusion Prompt Generation
- 引入傅里叶变换提取频域提示,结合空域提示实现多域特征学习
- 在三个主流数据集上达到领先性能,显著优于现有参数高效微调方法
- 适合需要高鲁棒性多模态视觉追踪的应用场景
近年来,视觉提示调优被引入RGB-热成像(RGB-T)追踪作为参数高效微调(PEFT)方法。然而,现有方法通常仅依赖空域信息作为提示,忽略了频域信息在提示学习中的关键作用。为此,我们提出一种高效视觉傅里叶提示追踪方法(VFPTrack),通过快速傅里叶变换(FFT)学习模态相关提示。该方法包含共享参数的对称特征提取编码器、视觉傅里叶提示模块,以及通过多模态特征融合生成双向交互提示的模态融合提示生成器。首先,使用冻结的特征提取编码器分别提取可见光(RGB)和热红外(TIR)模态特征;随后,将空域提示与FFT获得的频域提示相结合,充分挖掘不同域的信息。最后,不同于以往融合方式,本方法通过融合多模态特征生成联合模态提示,并与各单模态交互,实现跨模态充分特征交互。在三个主流RGB-T追踪基准上的大量实验表明,该方法表现优异。
原文摘要 · Abstract (English)
Recently, visual prompt tuning is introduced to RGB-Thermal (RGB-T) tracking as a parameter-efficient finetuning (PEFT) method. However, these PEFT-based RGB-T tracking methods typically rely solely on spatial domain information as prompts for feature extraction. As a result, they often fail to achieve optimal performance by overlooking the crucial role of frequency-domain information in prompt learning. To address this issue, we propose an efficient Visual Fourier Prompt Tracking (named VFPTrack) method to learn modality-related prompts via Fast Fourier Transform (FFT). Our method consists of symmetric feature extraction encoder with shared parameters, visual fourier prompts, and Modality Fusion Prompt Generator that generates bidirectional interaction prompts through multi-modal feature fusion. Specifically, we first use a frozen feature extraction encoder to extract RGB and thermal infrared (TIR) modality features. Then, we combine the visual prompts in the spatial domain with the frequency domain prompts obtained from the FFT, which allows for the full extraction and understanding of modality features from different domain information. Finally, unlike previous fusion methods, the modality fusion prompt generation module we use combines features from different modalities to generate a fused modality prompt. This modality prompt is interacted with each individual modality to fully enable feature interaction across different modalities. Extensive experiments conducted on three popular RGB-T tracking benchmarks show that our method demonstrates outstanding performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。