提出新系统Savvy与评估框架OGA,解决长视频中物体持续识别难题。
Open-World Video Segmentation

- 分层掩码发现+延迟准入+轨迹合并,实现长视频物体持续追踪。
- 在ScanNet和HM3D上超越基线,多项指标领先,尤其在身份持久性上表现突出。
- 引入粒度感知评估,更公平衡量开放世界视频分割性能,适合研究长期视觉理解者。
尽管视频分割在短片段和封闭集基准上进展迅速,但开放世界视频分割仍基本未被探索。挑战在于:(1) 现有方法不支持动态第一视角长视频中的物体发现与身份维持;(2) 评估协议采用严格的1:1匹配,不公平惩罚语义合理但粒度不一致的预测。为此,我们提出Savvy——一个面向零样本开放世界长时视频分割的实用且强大的系统。Savvy结合分层掩码发现、延迟准入与轨迹合并,支持持续物体发现、安全轨迹提升与稳定长程身份维持。同时提出OGA——一种粒度感知的评估套件。基于粒度无关(GA)匹配协议,将传统1:1匹配扩展为n:1映射,但仍通过断点检测保持时间严谨性,并通过主导连贯片段评分每个参考对象。这防止碎片化或闪烁支持被过度奖励,同时支持GA适配的指标与结构诊断:身份持久性(IP)与身份集中度(IC)。在VIPSeg上,标准1:1评估显著低估开放世界方法性能,而GA评估恢复了其被压制的表现。在更具现实性的长时基准(ScanNet和HM3D)上,Savvy在经典与新指标(包括STQ、VPQ$_\infty$、IP、IC)上均持续优于强基线。这些结果共同建立了开放世界长时视频分割的实用基准与强基线。
原文摘要 · Abstract (English)
While video segmentation has advanced rapidly on short clips and closed-set benchmarks, open-world video segmentation remains largely unexplored. The challenge is twofold: (1) existing methods are not designed to support object discovery and identity maintenance in long videos of dynamic ego-motion, and (2) existing evaluation protocols rely on a rigid 1:1 matching that unfairly penalizes semantically valid predictions with mismatched granularity. To address both gaps, we introduce Savvy, a practical and strong system for zero-shot open-world long-horizon video segmentation. Savvy combines hierarchical mask discovery, deferred admission, and track consolidation to support persistent object discovery, safe track promotion, and stable long-range identity maintenance. We further propose OGA, a granularity-aware evaluation suite for open-world video segmentation. Built on a Granularity-Agnostic (GA) matching protocol, OGA relaxes conventional 1:1 matching to an n:1 mapping, but still enforces temporal rigor by detecting support discontinuities through sever points and scoring each reference object through its dominant coherent fragment. This prevents fragmented or flickering support from being over-rewarded while enabling GA-adapted metrics and structural diagnostics: identity persistence (IP), and identity concentration (IC). On VIPSeg, we show that standard 1:1 evaluation substantially underestimates open-world methods, whereas GA evaluation recovers much of their suppressed performance. On the more realistic long-horizon benchmarks: ScanNet and HM3D, Savvy consistently outperforms strong baselines across both classical and proposed metrics, including STQ, VPQ$_\infty$, IP and IC. Together, these results establish a practical benchmark and a strong baseline for open-world long-horizon video segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。