arXiv:2506.19621cs.CVcs.AI2025-06中稿 · Publication at ICA…

用频域相位相关技术实现无监督视频对象分解与预测

VideoPCDNet: Video Parsing and Prediction with Phase Correlation Networks

  • 通过频域相位相关递归解析视频为对象组件
  • 在多个合成数据集上超越基线模型的追踪与预测性能
  • 可学习可解释的对象与运动表示,适合视觉理解研究

理解与预测视频内容对动态环境中的规划与推理至关重要。尽管已有进展,但无监督学习对象表征与动态仍具挑战。我们提出VideoPCDNet,一种用于对象中心视频分解与预测的无监督框架。该模型利用频域相位相关技术,递归地将视频解析为对象成分,这些成分以学习到的对象原型的变换形式表示,从而实现准确且可解释的追踪。通过结合频域操作与轻量级学习模块显式建模对象运动,VideoPCDNet实现了准确的无监督对象追踪与未来帧预测。实验表明,在多个合成数据集上,VideoPCDNet在无监督追踪与预测任务中优于多种对象中心基线模型,同时学习到可解释的对象与运动表征。

原文摘要 · Abstract (English)

Understanding and predicting video content is essential for planning and reasoning in dynamic environments. Despite advancements, unsupervised learning of object representations and dynamics remains challenging. We present VideoPCDNet, an unsupervised framework for object-centric video decomposition and prediction. Our model uses frequency-domain phase correlation techniques to recursively parse videos into object components, which are represented as transformed versions of learned object prototypes, enabling accurate and interpretable tracking. By explicitly modeling object motion through a combination of frequency domain operations and lightweight learned modules, VideoPCDNet enables accurate unsupervised object tracking and prediction of future video frames. In our experiments, we demonstrate that VideoPCDNet outperforms multiple object-centric baseline models for unsupervised tracking and prediction on several synthetic datasets, while learning interpretable object and motion representations.

视频生成无监督学习对象追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。