arXiv:2412.14295cs.CVcs.AI2024-12CVPR被引 34

提出对比槽时序一致性损失,提升视频物体中心表征的稳定性。

Temporally Consistent Object-Centric Learning by Contrasting Slots

  • 设计物体级时序对比损失,显式强化表征的时间一致性。
  • 在合成与真实数据集上实现当前最优物体发现效果。
  • 适合需要稳定物体追踪与动态预测的任务场景。

从视频中进行无监督物体中心学习是一种有前景的方法,可从大量未标注视频中提取结构化表征。为支持自主控制等下游任务,这些表征必须具备组合性与时间一致性。现有基于递归处理的方法常因训练目标不强制时序一致性而缺乏长期帧间稳定性。本文引入一种新型物体级时序对比损失,显式促进时序一致性。该方法显著提升了学习到的物体中心表征的时间一致性,得到更可靠的视频分解结果,有助于实现具有挑战性的无监督物体动态预测任务。此外,该损失带来的归纳偏置大幅改善了物体发现性能,在合成与真实数据集上均超越现有方法,甚至优于利用运动掩码作为额外线索的弱监督方法。

原文摘要 · Abstract (English)

Unsupervised object-centric learning from videos is a promising approach to extract structured representations from large, unlabeled collections of videos. To support downstream tasks like autonomous control, these representations must be both compositional and temporally consistent. Existing approaches based on recurrent processing often lack long-term stability across frames because their training objective does not enforce temporal consistency. In this work, we introduce a novel object-level temporal contrastive loss for video object-centric models that explicitly promotes temporal consistency. Our method significantly improves the temporal consistency of the learned object-centric representations, yielding more reliable video decompositions that facilitate challenging downstream tasks such as unsupervised object dynamics prediction. Furthermore, the inductive bias added by our loss strongly improves object discovery, leading to state-of-the-art results on both synthetic and real-world datasets, outperforming even weakly-supervised methods that leverage motion masks as additional cues.

物体中心时序一致性对比学习视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。