提出统一框架,解决红外缺失时多模态跟踪不稳问题
Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking

- 用时空条件去噪变换器融合短期与长期时序信息
- 在三个公开数据集上显著优于现有方法
- 适配有无红外的场景,无需改架构或参数
RGBT跟踪中模态缺失常导致特征表示不完整且不稳定,严重降低性能。现有方法通常尝试从可用模态恢复缺失模态,但在复杂场景下生成质量不佳。此外,当前方法在处理缺失与完整数据时灵活性有限。为此,我们提出时空条件去噪变换器(SCDT),通过整合空间线索与时序上下文,在统一框架内自适应地重建缺失模态并增强弱模态特征,实现鲁棒的模态缺失RGBT跟踪。SCDT利用近期历史帧的短期时序线索捕捉细粒度相关性,以及编码模态演化的长期时序线索捕捉全局上下文。通过联合利用长短时序上下文作为条件,逐步引导可用模态的噪声特征学习到可靠且时序一致的多模态表示。此外,SCDT引入噪声调制自适应机制,根据模态可用性动态调整行为,使单一框架可统一处理模态缺失与完整场景,无需改变架构或参数。在三个公开基准数据集上的大量实验表明,该方法持续优于现有最优方法。
原文摘要 · Abstract (English)
Missing modalities in RGBT tracking often lead to incomplete and unstable multimodal feature representations that greatly degrade the performance. Existing methods typically attempt to recover missing modalities from available ones, but the quality of data generated in challenging scenarios might be unsatisfactory. In addition, current approaches exhibit limited flexibility in processing both missing and complete data. To overcome these limitations, we propose a Spatio-temporal Conditional Denoising Transformer (SCDT), which integrates the spatial cues and the temporal context to adaptively perform information reconstruction of missing modalities and feature enhancement of weak modalities in a unified framework, for robust modality-missing RGBT tracking. In particular, SCDT leverages the short-term temporal cues from recent historical frames to capture the fine-grained temporal correlations and the long-term temporal cues encoding modality evolution to capture the global context. By jointly exploiting long short-term temporal contexts as the conditions, SCDT progressively guides noisy features of available modalities to learn reliable and temporally consistent multimodal representations. Furthermore, SCDT introduces a noisemodulated adaptation mechanism that dynamically adjusts its behavior according to the modal availability, enabling a single framework to unify feature learning under both modality-missing and complete scenarios without changing the architecture or parameters. Extensive experiments on three public benchmark datasets demonstrate that our method consistently outperforms state-of-the-art methods. The code is available here.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。