用自然语言描述追踪任意物体,解决通用目标追踪难题
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
- 结合视觉与运动信息,改进卡尔曼滤波实现更精准追踪
- 在复杂场景下对同类物体仍保持高追踪精度
- 适合需要灵活追踪未知物体的研究者和开发者
尽管多目标追踪(MOT)取得进展,仍依赖先验知识和预定义类别,难以追踪陌生物体。通用多目标追踪(GMOT)虽减少依赖,但现有方法多为一次性追踪(OneShot-GMOT),严重依赖初始边界框,在视角、光照、遮挡和尺度变化下表现不佳。为此,本文提出基于自然语言描述的地面化通用多目标追踪(Grounded-GMOT),用户可通过语言指定要追踪的物体属性。首先构建G2MOT数据集,包含多种通用物体及其详细属性文本描述。随后提出KAM-SORT方法,融合视觉外观与运动信息,并增强卡尔曼滤波。该方法在同类物体密集场景中表现优异。实验表明,Grounded-GMOT优于现有一次性追踪方法,且KAM-SORT在各类追踪器对比中展现出显著优势。
原文摘要 · Abstract (English)
Despite recent progress, Multi-Object Tracking (MOT) continues to face significant challenges, particularly its dependence on prior knowledge and predefined categories, complicating the tracking of unfamiliar objects. Generic Multiple Object Tracking (GMOT) emerges as a promising solution, requiring less prior information. Nevertheless, existing GMOT methods, primarily designed as OneShot-GMOT, rely heavily on initial bounding boxes and often struggle with variations in viewpoint, lighting, occlusion, and scale. To overcome the limitations inherent in both MOT and GMOT when it comes to tracking objects with specific generic attributes, we introduce Grounded-GMOT, an innovative tracking paradigm that enables users to track multiple generic objects in videos through natural language descriptors. Our contributions begin with the introduction of the G2MOT dataset, which includes a collection of videos featuring a wide variety of generic objects, each accompanied by detailed textual descriptions of their attributes. Following this, we propose a novel tracking method, KAM-SORT, which not only effectively integrates visual appearance with motion cues but also enhances the Kalman filter. KAM-SORT proves particularly advantageous when dealing with objects of high visual similarity from the same generic category in GMOT scenarios. Through comprehensive experiments, we demonstrate that Grounded-GMOT outperforms existing OneShot-GMOT approaches. Additionally, our extensive comparisons between various trackers highlight KAM-SORT's efficacy in GMOT, further establishing its significance in the field. Project page: https://UARK-AICV.github.io/G2MOT. The source code and dataset will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。