arXiv:2410.17534cs.CV2024-10NeurIPS被引 6

构建首个大规模开放词汇多目标追踪基准,推动视频目标追踪研究

OVT-B: A New Large-Scale Benchmark for Open-Vocabulary Multi-Object Tracking

  • 提出融合运动特征的追踪方法,弥补以往OVT忽略运动信息的缺陷
  • 构建包含1048类、1973段视频、63.7万标注的OVT-B数据集,规模超现有数据集5倍以上
  • 开源数据集与基线方法,助力开放词汇追踪领域发展

开放词汇目标感知已成为人工智能重要方向,旨在识别训练中未见的新类别物体。尽管单图像开放词汇检测已有广泛研究,但视频中的开放词汇多目标追踪(OVT)仍较少被关注,主要受限于缺乏基准数据集。本文构建了首个大规模开放词汇多目标追踪基准OVT-B,包含1,048个物体类别、1,973段视频及637,608个边界框标注,远超现有唯一OVT数据集OVTAO-val(200+类别,900+视频)。该基准可为OVT研究提供新标准。同时,我们提出一种简单有效的基线方法,融合运动特征进行追踪,而此前的OVT方法常忽略此关键信息。实验验证了基准的有效性及方法的优越性。数据集已公开于https://github.com/Coo1Sea/OVT-B-Dataset。

原文摘要 · Abstract (English)

Open-vocabulary object perception has become an important topic in artificial intelligence, which aims to identify objects with novel classes that have not been seen during training. Under this setting, open-vocabulary object detection (OVD) in a single image has been studied in many literature. However, open-vocabulary object tracking (OVT) from a video has been studied less, and one reason is the shortage of benchmarks. In this work, we have built a new large-scale benchmark for open-vocabulary multi-object tracking namely OVT-B. OVT-B contains 1,048 categories of objects and 1,973 videos with 637,608 bounding box annotations, which is much larger than the sole open-vocabulary tracking dataset, i.e., OVTAO-val dataset (200+ categories, 900+ videos). The proposed OVT-B can be used as a new benchmark to pave the way for OVT research. We also develop a simple yet effective baseline method for OVT. It integrates the motion features for object tracking, which is an important feature for MOT but is ignored in previous OVT methods. Experimental results have verified the usefulness of the proposed benchmark and the effectiveness of our method. We have released the benchmark to the public at https://github.com/Coo1Sea/OVT-B-Dataset.

多目标追踪开放词汇视频理解数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。