arXiv:2510.13235cs.CV2025-10

用显式隐式提示动态建模目标,提升多目标跟踪鲁棒性

EPIPTrack: Rethinking Prompt Modeling with Explicit and Implicit Prompts for Multi-Object Tracking

  • 设计显式(运动转文本)与隐式(可学习伪词)双提示机制
  • 在MOT17/MOT20/DanceTrack上均超越现有方法,性能显著提升
  • 适合需要语义理解与实时适应的复杂跟踪场景

多模态语义线索(如文本描述)在提升目标感知方面展现出强大潜力。然而,现有方法依赖大型语言模型生成的静态文本描述,难以适应目标状态的实时变化,且易产生幻觉。为此,我们提出统一的视觉-语言跟踪框架EPIPTrack,通过显式与隐式提示实现动态目标建模与语义对齐。显式提示将空间运动信息转化为自然语言描述,提供时空引导;隐式提示结合伪词与可学习描述符,构建个体化外观属性表征。两类提示均通过CLIP文本编码器动态调整,以响应目标状态变化。此外,设计判别性特征增强模块,强化视觉与跨模态表示。在MOT17、MOT20和DanceTrack上的大量实验表明,EPIPTrack在多样化场景中表现出更强的适应性与优越性能。

原文摘要 · Abstract (English)

Multimodal semantic cues, such as textual descriptions, have shown strong potential in enhancing target perception for tracking. However, existing methods rely on static textual descriptions from large language models, which lack adaptability to real-time target state changes and prone to hallucinations. To address these challenges, we propose a unified multimodal vision-language tracking framework, named EPIPTrack, which leverages explicit and implicit prompts for dynamic target modeling and semantic alignment. Specifically, explicit prompts transform spatial motion information into natural language descriptions to provide spatiotemporal guidance. Implicit prompts combine pseudo-words with learnable descriptors to construct individualized knowledge representations capturing appearance attributes. Both prompts undergo dynamic adjustment via the CLIP text encoder to respond to changes in target state. Furthermore, we design a Discriminative Feature Augmentor to enhance visual and cross-modal representations. Extensive experiments on MOT17, MOT20, and DanceTrack demonstrate that EPIPTrack outperforms existing trackers in diverse scenarios, exhibiting robust adaptability and superior performance.

多目标跟踪视觉语言模型动态提示语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。