arXiv:2603.05384cs.CV2026-03被引 1

提出全景视觉下的语言指代多目标追踪,解决传统方法视野受限问题。

ORMOT: A Dataset and Framework for Omnidirectional Referring Multi-Object Tracking

  • 构建全景视角下的语言指代追踪框架与数据集
  • 涵盖27个场景、848条语言描述、3401个标注目标
  • 适合研究全景视觉与多模态追踪的学者

多目标追踪(MOT)是计算机视觉中的基础任务,旨在跨视频帧追踪目标。现有MOT方法在常规视觉场景中表现良好,但在视觉-语言设置下面临显著挑战。为此,近期提出了语言指代多目标追踪(RMOT)任务,旨在追踪与语言描述对应的目标。然而,当前RMOT方法主要基于传统相机拍摄的数据集,存在视场有限的问题,导致目标频繁移出画面,造成追踪中断和上下文丢失。本文提出新任务——全景语言指代多目标追踪(ORMOT),将RMOT扩展至全景图像,以克服传统数据集的视场限制,并提升模型对长时语言描述的理解能力。为推动该任务发展,我们构建了ORSet数据集,包含27个多样化的全景场景、848条语言描述和3401个标注目标,提供丰富的视觉、时间和语言信息。同时,我们提出针对全景语言指代追踪的大型视觉-语言模型驱动框架ORTrack。在ORSet上的大量实验验证了该框架的有效性。数据集与代码将在https://github.com/chen-si-jia/ORMOT公开。

原文摘要 · Abstract (English)

Multi-Object Tracking (MOT) is a fundamental task in computer vision, aiming to track targets across video frames. Existing MOT methods perform well in general visual scenes, but face significant challenges and limitations when extended to visual-language settings. To bridge this gap, the task of Referring Multi-Object Tracking (RMOT) has recently been proposed, which aims to track objects that correspond to language descriptions. However, current RMOT methods are primarily developed on datasets captured by conventional cameras, which suffer from limited field of view. This constraint often causes targets to move out of the frame, leading to fragmented tracking and loss of contextual information. In this work, we propose a novel task, called Omnidirectional Referring Multi-Object Tracking (ORMOT), which extends RMOT to omnidirectional imagery, aiming to overcome the field-of-view (FoV) limitation of conventional datasets and improve the model's ability to understand long-horizon language descriptions. To advance the ORMOT task, we construct ORSet, an Omnidirectional Referring Multi-Object Tracking dataset, which contains 27 diverse omnidirectional scenes, 848 language descriptions, and 3,401 annotated objects, providing rich visual, temporal, and language information. Furthermore, we propose ORTrack, a Large Vision-Language Model (LVLM)-driven framework tailored for Omnidirectional Referring Multi-Object Tracking. Extensive experiments on the ORSet dataset demonstrate the effectiveness of our ORTrack framework. The dataset and code will be open-sourced at https://github.com/chen-si-jia/ORMOT.

多目标追踪全景视觉语言指代视觉-语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。