arXiv:2508.06335cs.CV2025-08

无监督视频预测中,模型能从观测中自动推断状态,不再依赖初始真值。

ViPro-2: Unsupervised State Estimation via Integrated Dynamics for Guiding Video Prediction

  • 通过整合动态模型,实现无需初始真值的状态推断
  • 在3D扩展数据集上验证了无监督状态估计的有效性
  • 适合研究视频生成与自监督学习的学者

视频未来帧预测是具有广泛应用价值的挑战性任务。先前工作表明,过程性知识可帮助深度模型处理复杂动态场景,但原模型ViPro依赖给定的初始符号状态。我们发现该设定导致模型学习到不合理的捷径,当先前状态存在噪声时无法正确估计当前状态。本文对ViPro进行多项改进,使模型能在无监督条件下仅凭观测准确推断状态,无需提供完整初始真值。我们在原Orbits数据集基础上扩展出3D版本,更贴近真实世界场景,验证了方法的有效性。

原文摘要 · Abstract (English)

Predicting future video frames is a challenging task with many downstream applications. Previous work has shown that procedural knowledge enables deep models for complex dynamical settings, however their model ViPro assumed a given ground truth initial symbolic state. We show that this approach led to the model learning a shortcut that does not actually connect the observed environment with the predicted symbolic state, resulting in the inability to estimate states given an observation if previous states are noisy. In this work, we add several improvements to ViPro that enables the model to correctly infer states from observations without providing a full ground truth state in the beginning. We show that this is possible in an unsupervised manner, and extend the original Orbits dataset with a 3D variant to close the gap to real world scenarios.

视频预测无监督学习状态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。