arXiv:2606.29861cs.CVcs.AI2026-06

用非线性状态空间模型统一追踪与分割运动物体

SUMO: Segment and Track Any Motion with Nonlinear State Space Models

论文配图:SUMO: Segment and Track Any Motion with Nonlinear State Space Models
图 1 · 摘自论文原文
  • 基于机器人学原理设计非线性状态空间模型捕捉复杂运动
  • 提出选择性无迹滤波器,动态融合多源预测提升状态估计精度
  • 零样本训练、无需微调,适合复杂运动场景下的视觉追踪

视觉目标追踪(VOT)与运动目标分割(MOS)是计算机视觉中涉及时空动态的两个基础任务。现有方法主要依赖视觉线索,在物体运动复杂且非线性的现实场景中表现不佳。为此,我们提出SUMO,一种零样本、无需训练的统一框架,将非线性动力学与基于视觉的分割相结合,实现准确一致的VOT与MOS。具体而言,我们设计了一种受机器人学启发的非线性状态空间模型(SSM),以捕捉复杂的物体动态;在此基础上,提出选择性无迹滤波器(SUF),通过联合评分机制和动态融合多源预测,识别随时间变化的最可能物体状态;此外,引入记忆选择机制评估记忆帧的可靠性。大量实验表明,SUMO在VOT与MOS任务上均达到领先性能。

原文摘要 · Abstract (English)

Visual Object Tracking (VOT) and Moving Object Segmentation (MOS) are two fundamental tasks in computer vision that involve both spatial and temporal object dynamics. Existing methods rely predominantly on visual cues and thus often falter in real-world scenarios where object motions are inherently complex and nonlinear. To address this limitation, we propose SUMO, a zero-shot, training-free, unified framework integrating nonlinear dynamics with vision-based segmentation for accurate and consistent VOT and MOS. Specifically, we develop a nonlinear State Space Model (SSM) inspired by robotics principles to capture the complex object dynamics. Building on this model, we propose a Selective Unscented Filter (SUF) for accurate state estimation, which features a joint scoring mechanism and dynamically fuses multi-source predictions to identify the most plausible object state over time. Furthermore, we apply a memory selection mechanism to evaluate the reliability of memory frames. Our extensive experimental results show that SUMO achieves state-of-the-art performance on both VOT and MOS tasks.

目标追踪运动分割状态空间模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。