综述跨模态视觉目标跟踪最新进展,涵盖多传感器融合方法。
Visual Object Tracking across Diverse Data Modalities: A Review
- 分类梳理单模态(RGB/热成像/点云)与多模态(如RGB-热、RGB-LiDAR)跟踪框架
- 总结四大主流多模态组合在多个基准上的性能对比结果
- 适合关注多传感器目标跟踪与深度学习应用的研究者阅读
视觉目标跟踪(VOT)是计算机视觉中一个关键且具有吸引力的研究方向,旨在识别并追踪视频序列中任意类别的目标。该技术可应用于多种数据模态场景,包括可见光(RGB)、热红外和点云等。由于单一传感器难以应对复杂多变的环境,多模态视觉目标跟踪也受到广泛关注。本文全面综述了单模态与多模态VOT的最新进展,尤其聚焦深度学习方法。首先,系统回顾三类主流单模态追踪方法:基于RGB、热红外与点云的跟踪;归纳出四种广泛使用的单模态框架,抽象其结构并分类现有模型。其次,总结四类多模态追踪:RGB-深度、RGB-热红外、RGB-LiDAR与RGB-语言。此外,呈现多种基准上各类模态的对比实验结果。最后,提出未来研究建议与深刻见解,推动该快速发展的领域持续演进。
原文摘要 · Abstract (English)
Visual Object Tracking (VOT) is an attractive and significant research area in computer vision, which aims to recognize and track specific targets in video sequences where the target objects are arbitrary and class-agnostic. The VOT technology could be applied in various scenarios, processing data of diverse modalities such as RGB, thermal infrared and point cloud. Besides, since no one sensor could handle all the dynamic and varying environments, multi-modal VOT is also investigated. This paper presents a comprehensive survey of the recent progress of both single-modal and multi-modal VOT, especially the deep learning methods. Specifically, we first review three types of mainstream single-modal VOT, including RGB, thermal infrared and point cloud tracking. In particular, we conclude four widely-used single-modal frameworks, abstracting their schemas and categorizing the existing inheritors. Then we summarize four kinds of multi-modal VOT, including RGB-Depth, RGB-Thermal, RGB-LiDAR and RGB-Language. Moreover, the comparison results in plenty of VOT benchmarks of the discussed modalities are presented. Finally, we provide recommendations and insightful observations, inspiring the future development of this fast-growing literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。