用频域分析提升视觉事件追踪,解决高速低光场景跟踪难题
FreqTrack: Frequency Learning based Vision Transformer for RGB-Event Object Tracking

- 在频域而非空间域融合RGB与事件数据,利用傅里叶滤波动态增强特征
- 在COESOT数据集上达到76.6%精度,领先现有方法,尤其在高速低光下表现优异
- 适合需要高鲁棒性视觉追踪的自动驾驶、机器人等实时应用
现有单模态RGB追踪器在复杂动态场景中性能受限,而事件传感器为提升追踪能力带来新可能。然而,当前多数RGB-事件融合方法主要基于卷积、Transformer或Mamba架构,在空间域进行融合,未能充分挖掘事件数据的独特时间响应和高频特性。为此,我们提出FreqTrack,一种基于频域建模的RGBE追踪框架,通过频域变换建立跨模态互补关联,实现更鲁棒的特征融合。设计了谱增强Transformer(SET)层,引入多头动态傅里叶滤波,自适应增强与选择频域特征;同时开发小波边缘精炼(WER)模块,利用可学习小波变换显式提取事件数据的多尺度边缘结构,显著提升高速与低光场景下的建模能力。在COESOT和FE108数据集上的大量实验表明,FreqTrack表现极具竞争力,尤其在COESOT基准上取得76.6%的领先精度,验证了频域建模在RGBE追踪中的有效性。
原文摘要 · Abstract (English)
Existing single-modal RGB trackers often face performance bottlenecks in complex dynamic scenes, while the introduction of event sensors offers new potential for enhancing tracking capabilities. However, most current RGB-event fusion methods, primarily designed in the spatial domain using convolutional, Transformer, or Mamba architectures, fail to fully exploit the unique temporal response and high-frequency characteristics of event data. To address this, we1 propose FreqTrack, a frequency-aware RGBE tracking framework that establishes complementary inter-modal correlations through frequency-domain transformations for more robust feature fusion. We design a Spectral Enhancement Transformer (SET) layer that incorporates multi-head dynamic Fourier filtering to adaptively enhance and select frequency-domain features. Additionally, we develop a Wavelet Edge Refinement (WER) module, which leverages learnable wavelet transforms to explicitly extract multi-scale edge structures from event data, effectively improving modeling capability in high-speed and low-light scenarios. Extensive experiments on the COESOT and FE108 datasets demonstrate that FreqTrack achieves highly competitive performance, particularly attaining leading precision of 76.6\% on the COESOT benchmark, validating the effectiveness of frequency-domain modeling for RGBE tracking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。