提出3D时空感知的运动物体分割框架,实现精准动态追踪
GMOS: Grounding Moving Object Segmentation in 3D Space and Time

- 直接处理视频输入,融合3D空间与时间信息进行分割
- 在5个基准上达领先性能,推理速度显著快于现有方法
- 适合需要实时流式部署的多目标运动追踪场景
运动物体分割(MOS)旨在发现、分割并追踪独立于相机运动的物体。现有方法存在两大局限:依赖缺乏3D几何信息的预计算2D辅助模态(如光流或点轨迹),且将运动视为序列级属性,忽视每个物体的瞬时运动状态。为此,本文提出GMOS框架,直接基于RGB视频实现3D感知、时间精细的多物体运动分割,并提供快速部署版本GMOS-S。为支持该范式训练与评估,构建了包含2,210个真实世界视频的GMOS-2K数据集,涵盖五个主流视频对象分割(VOS)基准的逐物体时间运动标注,并提出MOS-I(I代表瞬时)评估协议,包含三个互补指标。GMOS在MOS、MOS-I及无监督VOS基准上均达当前最优表现,同时推理速度显著优于先前多物体MOS方法,支持在线推理用于流式部署。
原文摘要 · Abstract (English)
Moving Object Segmentation (MOS) aims to discover, segment, and track objects that move independently of the camera. Current MOS methods, however, exhibit two fundamental limitations: they rely on pre-computed 2D auxiliary modalities such as optical flow or point trajectories that lack 3D geometric information, and they treat motion as a sequence-level attribute, overlooking the instantaneous motion state of each object. We address both by grounding MOS in 3D space and time, and propose GMOS, a framework that operates directly on RGB video to produce 3D-aware, temporally fine-grained segmentation of multiple moving objects, alongside a foreground--background variant GMOS-S for faster deployment. To support training and evaluation in this regime, we curate GMOS-2K, a dataset of 2,210 real-world videos with per-object temporal motion annotations drawn from five established Video Object Segmentation (VOS) benchmarks, and formalise MOS-I ("I" for instantaneous), a temporally fine-grained evaluation protocol with three complementary metrics. GMOS achieves state-of-the-art results across MOS, MOS-I, and Unsupervised VOS benchmarks, while running significantly faster than prior multi-object MOS methods and supporting online inference for streaming deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。