arXiv:2511.19475cs.CVcs.AI2025-11AAAI被引 3

统一多模态视频目标跟踪与分割,提升模型泛化能力。

Tracking and Segmenting Anything in Any Modality

  • 采用解耦专家混合机制,分离跨模态共性与任务特异性特征。
  • 在18个基准上表现优异,显著提升多任务联合训练效果。
  • 适合需要跨模态、多任务通用模型的研究与应用者。

跟踪与分割在视频理解中至关重要,提供物体在视频序列中的位置信息与时间关联。尽管目标一致,现有方法通常使用专用架构或模态特定参数,限制了泛化与可扩展性。近期工作尝试从任意模态输入或多任务推理角度统一多个子任务,但常忽视不同模态间的分布差异和任务间特征表示差异,阻碍跨任务与跨模态知识共享,制约通用模型发展。为此,我们提出统一的跟踪与分割框架SATA,支持任意模态输入下的广泛子任务统一。具体地,引入解耦专家混合(DeMoE)机制,将统一表征学习解耦为跨模态共享知识建模与特定信息建模过程,增强灵活性与泛化能力。同时,设计任务感知多目标跟踪(TaMOT)流水线,将所有任务输出统一为带校准ID实例集合,缓解多任务训练中任务特异性知识退化问题。SATA在18个挑战性基准上表现卓越,为更通用的视频理解提供了新视角。

原文摘要 · Abstract (English)

Tracking and segmentation play essential roles in video understanding, providing basic positional information and temporal association of objects within video sequences. Despite their shared objective, existing approaches often tackle these tasks using specialized architectures or modality-specific parameters, limiting their generalization and scalability. Recent efforts have attempted to unify multiple tracking and segmentation subtasks from the perspectives of any modality input or multi-task inference. However, these approaches tend to overlook two critical challenges: the distributional gap across different modalities and the feature representation gap across tasks. These issues hinder effective cross-task and cross-modal knowledge sharing, ultimately constraining the development of a true generalist model. To address these limitations, we propose a universal tracking and segmentation framework named SATA, which unifies a broad spectrum of tracking and segmentation subtasks with any modality input. Specifically, a Decoupled Mixture-of-Expert (DeMoE) mechanism is presented to decouple the unified representation learning task into the modeling process of cross-modal shared knowledge and specific information, thus enabling the model to maintain flexibility while enhancing generalization. Additionally, we introduce a Task-aware Multi-object Tracking (TaMOT) pipeline to unify all the task outputs as a unified set of instances with calibrated ID information, thereby alleviating the degradation of task-specific knowledge during multi-task training. SATA demonstrates superior performance on 18 challenging tracking and segmentation benchmarks, offering a novel perspective for more generalizable video understanding.

视频理解多模态跟踪分割通用模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。