arXiv:2412.17807cs.CVcs.AI2024-12AAAI被引 16

跨视角语言引导多目标跟踪,解决单视角遮挡问题。

Cross-View Referring Multi-Object Tracking

  • 引入多视角信息提升目标可见性,实现更准确的语义匹配。
  • 构建包含13场景、221条描述的CRTrack数据集,推动新任务发展。
  • 提出端到端的CRTracker方法,适用于复杂场景下的多目标追踪。

referring multi-object tracking (RMOT) 是当前追踪领域的重要课题,其任务是根据语言描述引导追踪器定位并跟踪匹配的对象。现有研究主要聚焦于单视角下的RMOT,即单一视角序列或多组无关视角序列。然而,在单视角下,部分目标外观易被遮挡,导致与语言描述匹配错误。为此,本文提出新任务——跨视角语言引导多目标跟踪(CRMOT),通过融合多视角信息获取目标完整外观,避免单视角下因遮挡引起的误匹配。CRMOT要求在跨视角中精准追踪匹配语言描述的目标,并保持身份一致性。为推动该任务发展,我们基于CAMPUS和DIVOTrack数据集构建了跨视角语言引导多目标跟踪基准数据集CRTrack,包含13个不同场景和221条语言描述。同时,提出一种端到端的跨视角追踪方法CRTracker。在CRTrack上的大量实验验证了该方法的有效性。相关数据集与代码已公开于https://github.com/chen-si-jia/CRMOT。

原文摘要 · Abstract (English)

Referring Multi-Object Tracking (RMOT) is an important topic in the current tracking field. Its task form is to guide the tracker to track objects that match the language description. Current research mainly focuses on referring multi-object tracking under single-view, which refers to a view sequence or multiple unrelated view sequences. However, in the single-view, some appearances of objects are easily invisible, resulting in incorrect matching of objects with the language description. In this work, we propose a new task, called Cross-view Referring Multi-Object Tracking (CRMOT). It introduces the cross-view to obtain the appearances of objects from multiple views, avoiding the problem of the invisible appearances of objects in RMOT task. CRMOT is a more challenging task of accurately tracking the objects that match the language description and maintaining the identity consistency of objects in each cross-view. To advance CRMOT task, we construct a cross-view referring multi-object tracking benchmark based on CAMPUS and DIVOTrack datasets, named CRTrack. Specifically, it provides 13 different scenes and 221 language descriptions. Furthermore, we propose an end-to-end cross-view referring multi-object tracking method, named CRTracker. Extensive experiments on the CRTrack benchmark verify the effectiveness of our method. The dataset and code are available at https://github.com/chen-si-jia/CRMOT.

多目标追踪跨视角语言引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。