用视觉语言模型实现无需训练的多对象跟踪定位,精准响应自然语言查询。
Multi-Object Tracking Retrieval with LLaVA-Video: A Training-Free Solution to MOT25-StAG Challenge
- 将多对象跟踪转为视频检索任务,结合FastTracker与LLaVA-Video零样本推理。
- 在MOT25-StAG测试集上达到m-HIoU 20.68、HOTA 10.73,获挑战赛第二名。
- 无需训练即可处理自由语言查询,适合复杂场景下的实时多目标定位应用。
本文提出针对MOT25-Spatiotemporal Action Grounding(MOT25-StAG)挑战的解决方案。该挑战要求在复杂真实场景视频中,准确地定位并跟踪与特定自由语言查询匹配的多个对象。我们将其建模为视频检索问题,提出一种两阶段、零样本方法,融合SOTA追踪模型FastTracker与多模态大语言模型LLaVA-Video的优势。在MOT25-StAG测试集上,该方法取得m-HIoU 20.68和HOTA 10.73的性能,荣获挑战赛第二名。
原文摘要 · Abstract (English)
In this report, we present our solution to the MOT25-Spatiotemporal Action Grounding (MOT25-StAG) Challenge. The aim of this challenge is to accurately localize and track multiple objects that match specific and free-form language queries, using video data of complex real-world scenes as input. We model the underlying task as a video retrieval problem and present a two-stage, zero-shot approach, combining the advantages of the SOTA tracking model FastTracker and Multi-modal Large Language Model LLaVA-Video. On the MOT25-StAG test set, our method achieves m-HIoU and HOTA scores of 20.68 and 10.73 respectively, which won second place in the challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。