提出TenRMOT模型,实现语言引导的多目标视频跟踪与分割。
Temporal-Enhanced Multimodal Transformer for Referring Multi-Object Tracking and Segmentation
- 分阶段融合视觉与语言特征,增强跨模态理解。
- 引入时序更新模块,提升目标轨迹一致性。
- 构建新数据集Ref-KITTI Segmentation,支持多掩码标注。
指称多目标跟踪(RMOT)是一项新兴的跨模态任务,旨在通过语言表达定位任意数量的目标并维持其身份。该任务涉及语言与视觉模态的推理及目标的时间关联。现有方法仅采用松散特征融合,忽视了长期追踪信息。本文提出一种紧凑的Transformer方法TenRMOT,分别在编码和解码阶段进行特征融合。编码阶段逐层执行跨模态融合;解码阶段使用语言引导查询探测记忆特征以精准预测目标。此外,设计查询更新模块,显式利用目标的时序先验信息,增强轨迹一致性。同时提出新任务“指称多目标跟踪与分割”(RMOTS),构建新数据集Ref-KITTI Segmentation,包含18段视频、818个表达,平均每表达10.7个掩码,挑战远超多数现有单掩码数据集。TenRMOT在跟踪与分割任务上均表现优异。
原文摘要 · Abstract (English)
Referring multi-object tracking (RMOT) is an emerging cross-modal task that aims to locate an arbitrary number of target objects and maintain their identities referred by a language expression in a video. This intricate task involves the reasoning of linguistic and visual modalities, along with the temporal association of target objects. However, the seminal work employs only loose feature fusion and overlooks the utilization of long-term information on tracked objects. In this study, we introduce a compact Transformer-based method, termed TenRMOT. We conduct feature fusion at both encoding and decoding stages to fully exploit the advantages of Transformer architecture. Specifically, we incrementally perform cross-modal fusion layer-by-layer during the encoding phase. In the decoding phase, we utilize language-guided queries to probe memory features for accurate prediction of the desired objects. Moreover, we introduce a query update module that explicitly leverages temporal prior information of the tracked objects to enhance the consistency of their trajectories. In addition, we introduce a novel task called Referring Multi-Object Tracking and Segmentation (RMOTS) and construct a new dataset named Ref-KITTI Segmentation. Our dataset consists of 18 videos with 818 expressions, and each expression averages 10.7 masks, which poses a greater challenge compared to the typical single mask in most existing referring video segmentation datasets. TenRMOT demonstrates superior performance on both the referring multi-object tracking and the segmentation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。