系统梳理多模态目标跟踪的六大任务与核心挑战
Omni Survey for Multimodality Analysis in Visual Object Tracking
- 从数据、对齐、模型设计到评估,全面分析多模态跟踪四大差异
- 涵盖338篇文献,归纳不同模态组合的融合策略与实验配置
- 首次揭示现有数据集动物类别缺失与长尾分布问题,适合研究者参考
智慧城市发展催生了海量多模态数据,推动了多模态视觉目标跟踪(MMVOT)的关键进展。本文从多模态分析视角,系统梳理该任务的四大核心差异:数据采集、模态对齐与标注、模型设计及评估。首先介绍红外(T)、深度(D)、事件(E)、近红外(NIR)、语言(L)、声纳(S)等多模态数据特性,进而探讨其采集、对齐与标注难点。随后,基于对可见光(RGB)与X模态的处理方式,将现有方法分为复制/非复制分支的实验配置。最后讨论评估与基准测试。全文涵盖6类MMVOT任务,引用338篇文献,并首次分析现有数据集对象类别分布,发现其显著长尾特征及动物类别严重缺失,相较于纯RGB数据集。同时提出根本性问题:多模态融合是否总优于单模态?在何种条件下有效?
原文摘要 · Abstract (English)
The development of smart cities has led to the generation of massive amounts of multi-modal data in the context of a range of tasks that enable a comprehensive monitoring of the smart city infrastructure and services. This paper surveys one of the most critical tasks, multi-modal visual object tracking (MMVOT), from the perspective of multimodality analysis. Generally, MMVOT differs from single-modal tracking in four key aspects, data collection, modality alignment and annotation, model designing, and evaluation. Accordingly, we begin with an introduction to the relevant data modalities, laying the groundwork for their integration. This naturally leads to a discussion of challenges of multi-modal data collection, alignment, and annotation. Subsequently, existing MMVOT methods are categorised, based on different ways to deal with visible (RGB) and X modalities: programming the auxiliary X branch with replicated or non-replicated experimental configurations from the RGB branch. Here X can be thermal infrared (T), depth (D), event (E), near infrared (NIR), language (L), or sonar (S). The final part of the paper addresses evaluation and benchmarking. In summary, we undertake an omni survey of all aspects of multi-modal visual object tracking (VOT), covering six MMVOT tasks and featuring 338 references in total. In addition, we discuss the fundamental rhetorical question: Is multi-modal tracking always guaranteed to provide a superior solution to unimodal tracking with the help of information fusion, and if not, in what circumstances its application is beneficial. Furthermore, for the first time in this field, we analyse the distributions of the object categories in the existing MMVOT datasets, revealing their pronounced long-tail nature and a noticeable lack of animal categories when compared with RGB datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。