用确定性匹配替代预测,零参数实现视频对象一致性
Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence

- 用匈牙利算法在帧间做确定性匹配,取代传统预测模块
- 在MOVi-D/E和YouTube-VIS上表现媲美有参数方法
- 适合追求轻量高效视频对象建模的研究者
视频对象中心学习的主流方法通过可学习的动力学模块预测未来对象表示(槽),即所谓的槽。我们证明这些预测器本质上是离散对应问题的昂贵近似。现代自监督视觉主干已编码出可靠的实例区分特征。利用这些特征可省去学习的时间预测。我们提出Grounded Correspondence框架,将学习的转换函数替换为确定性的二分图匹配。槽从冻结主干特征中的显著区域初始化,帧间身份通过槽表示的匈牙利匹配保持。该方法在时间建模上无需可学习参数,但在MOVi-D、MOVi-E和YouTube-VIS上仍达到竞争力的表现。
原文摘要 · Abstract (English)
The de facto approach in video object-centric learning maintains temporal consistency through learned dynamics modules that predict future object representations, called slots. We demonstrate that these predictors function as expensive approximations of discrete correspondence problems. Modern self-supervised vision backbones already encode instance-discriminative features that distinguish objects reliably. Exploiting these features eliminates the need for learned temporal prediction. We introduce Grounded Correspondence, a framework that replaces learned transition functions with deterministic bipartite matching. Slots initialize from salient regions in frozen backbone features. Frame-to-frame identity is maintained through Hungarian matching on slot representations. The approach requires zero learnable parameters for temporal modeling yet achieves competitive performance on MOVi-D, MOVi-E, and YouTube-VIS. Project page: https://magenta-sherbet-85b101.netlify.app/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。