arXiv:2511.01427cs.CVcs.AI2025-11TPAMI被引 6

统一框架实现多模态目标追踪,支持多种输入与参考方式。

UniSOT: A Unified Framework for Multi-Modality Single Object Tracking

论文配图:UniSOT: A Unified Framework for Multi-Modality Single Object Tracking
图 1 · 摘自论文原文
  • 设计统一网络结构,兼容三种参考模态和四种视频模态
  • 在18个基准上超越专用追踪器,跨模态表现提升超3.0% AUC
  • 适合需要灵活适配多场景的工业级追踪应用

单目标追踪旨在通过特定参考模态(边界框、自然语言或两者)在特定视频模态(RGB、RGB+Depth、RGB+Thermal 或 RGB+Event)序列中定位目标。不同参考模态支持多样人机交互,不同视频模态则增强复杂场景下的鲁棒性。现有追踪器通常仅针对单一或有限组合的视频与参考模态设计,导致模型分散,难以实际部署。目前尚无追踪器能同时处理上述三类参考模态与四类视频模态。为此,本文提出统一追踪框架UniSOT,采用统一参数处理所有组合。在18个视觉追踪、视觉-语言追踪及RGB+X追踪基准上的实验表明,UniSOT性能显著优于各类专用追踪器。尤其在TNL2K数据集上,跨三种参考模态均提升超3.0% AUC;在所有RGB+X模态上,主指标超越Un-Track超2.0%。

原文摘要 · Abstract (English)

Single object tracking aims to localize target object with specific reference modalities (bounding box, natural language or both) in a sequence of specific video modalities (RGB, RGB+Depth, RGB+Thermal or RGB+Event.). Different reference modalities enable various human-machine interactions, and different video modalities are demanded in complex scenarios to enhance tracking robustness. Existing trackers are designed for single or several video modalities with single or several reference modalities, which leads to separate model designs and limits practical applications. Practically, a unified tracker is needed to handle various requirements. To the best of our knowledge, there is still no tracker that can perform tracking with these above reference modalities across these video modalities simultaneously. Thus, in this paper, we present a unified tracker, UniSOT, for different combinations of three reference modalities and four video modalities with uniform parameters. Extensive experimental results on 18 visual tracking, vision-language tracking and RGB+X tracking benchmarks demonstrate that UniSOT shows superior performance against modality-specific counterparts. Notably, UniSOT outperforms previous counterparts by over 3.0\% AUC on TNL2K across all three reference modalities and outperforms Un-Track by over 2.0\% main metric across all three RGB+X video modalities.

目标追踪多模态统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。