arXiv:2412.12561cs.CVcs.AI2024-12被引 3

通过语言引导提升多目标跟踪中新生目标的检测能力

Tell Me What to Track: Infusing Robust Language Guidance for Enhanced Referring Multi-Object Tracking

  • 设计协同匹配策略缓解新旧目标数据分布不均问题
  • 在编码器融合多模态特征,解码器注入语言查询引导
  • 适合需要精准语言驱动跟踪的应用场景

指称多目标跟踪(RMOT)是一项新兴的跨模态任务,旨在根据语言描述定位任意数量的目标,并在视频中持续跟踪。该任务需处理多模态数据并实现精确的目标定位与时间关联。然而,现有方法忽视了因任务特性导致的新生目标与已有目标间的数据分布不平衡问题,且仅间接融合多模态特征,难以明确指导新生目标检测。为此,本文提出协同匹配策略,缓解不平衡影响,增强新生目标检测能力,同时保持跟踪性能。在编码器中,整合并强化跨模态与多尺度特征融合,克服以往工作中多模态信息共享与交互受限的瓶颈。在解码器中,引入指称感知适配机制,通过查询令牌提供显式语言引导。实验表明,所提模型相比先前方法性能提升3.42%,验证了设计的有效性。

原文摘要 · Abstract (English)

Referring multi-object tracking (RMOT) is an emerging cross-modal task that aims to localize an arbitrary number of targets based on a language expression and continuously track them in a video. This intricate task involves reasoning on multi-modal data and precise target localization with temporal association. However, prior studies overlook the imbalanced data distribution between newborn targets and existing targets due to the nature of the task. In addition, they only indirectly fuse multi-modal features, struggling to deliver clear guidance on newborn target detection. To solve the above issues, we conduct a collaborative matching strategy to alleviate the impact of the imbalance, boosting the ability to detect newborn targets while maintaining tracking performance. In the encoder, we integrate and enhance the cross-modal and multi-scale fusion, overcoming the bottlenecks in previous work, where limited multi-modal information is shared and interacted between feature maps. In the decoder, we also develop a referring-infused adaptation that provides explicit referring guidance through the query tokens. The experiments showcase the superior performance of our model (+3.42%) compared to prior works, demonstrating the effectiveness of our designs.

多目标跟踪语言引导跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。