arXiv:2410.08529cs.CVcs.AI2024-10被引 5

提出视频驱动的开放词汇目标追踪方法,提升对新类别物体的追踪能力。

VOVTrack: Exploring the Potentiality in Videos for Open-Vocabulary Object Tracking

  • 基于视频时序特性设计提示引导注意力机制,提升动态物体定位精度。
  • 利用无标注视频数据自监督学习物体相似性,实现跨帧关联追踪。
  • 适合需要追踪未见类别的实际视频场景,如自动驾驶与监控系统。

开放词汇多目标追踪(OVMOT)是视频中检测与追踪多样物体类别的一项关键挑战,涵盖已见类别(基础类)和未见类别(新类)。该任务融合了开放词汇目标检测(OVD)与多目标追踪(MOT)的复杂性。现有方法通常将两者作为独立模块拼接,且以图像为中心。本文提出VOVTrack,从视频对象追踪视角出发,整合运动状态信息与视频级训练策略。首先,引入追踪相关的目标状态建模,并设计提示引导注意力机制,提升对随时间变化物体的定位与分类精度;其次,通过自监督物体相似性学习,在无标注原始视频上进行训练,实现时序上的物体关联。实验表明,VOVTrack显著优于现有方法,成为开放词汇追踪任务的当前最优方案。

原文摘要 · Abstract (English)

Open-vocabulary multi-object tracking (OVMOT) represents a critical new challenge involving the detection and tracking of diverse object categories in videos, encompassing both seen categories (base classes) and unseen categories (novel classes). This issue amalgamates the complexities of open-vocabulary object detection (OVD) and multi-object tracking (MOT). Existing approaches to OVMOT often merge OVD and MOT methodologies as separate modules, predominantly focusing on the problem through an image-centric lens. In this paper, we propose VOVTrack, a novel method that integrates object states relevant to MOT and video-centric training to address this challenge from a video object tracking standpoint. First, we consider the tracking-related state of the objects during tracking and propose a new prompt-guided attention mechanism for more accurate localization and classification (detection) of the time-varying objects. Subsequently, we leverage raw video data without annotations for training by formulating a self-supervised object similarity learning technique to facilitate temporal object association (tracking). Experimental results underscore that VOVTrack outperforms existing methods, establishing itself as a state-of-the-art solution for open-vocabulary tracking task.

目标追踪开放词汇视频理解自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。