用预测未来帧提升视频生成质量,让潜空间更懂运动规律。
Video Generation with Predictive Latents

- 通过随机丢弃未来帧,让模型同时重建过去和预测未来。
- 在UCF101上训练速度加快52%,生成质量提升34.42点FVD。
- 适合追求高效生成与运动一致性的视频建模研究者。
视频变分自编码器(VAE)通过将视觉世界映射到紧凑的时空潜空间,实现高效的视频生成建模。尽管现有方法在重建质量上表现良好,但持续优化重建并不必然提升生成性能。如何增强视频潜变量的可扩散性仍是关键挑战。受预测世界建模原理启发,本文提出一种统一预测学习与视频重建的简单有效目标:随机丢弃未来帧,仅用部分历史观测进行编码,训练解码器同时重建已知帧并预测未来帧。该设计促使潜空间学习时间预测结构,建立对视频动态的更连贯理解,从而提升生成质量。所提模型名为预测式视频VAE(PV-VAE),在视频生成任务中表现优异,在UCF101数据集上相比Wan2.2 VAE实现52%的收敛加速与34.42点的FVD提升。综合分析表明,PV-VAE具备良好可扩展性,生成性能随训练提升,且在下游视频理解任务中持续获益,证明其潜空间有效捕捉了时间一致性与运动先验。
原文摘要 · Abstract (English)
Video Variational Autoencoder (VAE) enables latent video generative modeling by mapping the visual world into compact spatiotemporal latent spaces, improving training efficiency and stability. While existing video VAEs achieve commendable reconstruction quality, continued optimization of reconstruction does not necessarily translate into improved generative performance. How to enhance the diffusability of video latents remains a critical and unresolved challenge. In this work, inspired by principles of predictive world modeling, we investigate the potential of predictive learning to improve the video generative modeling. To this end, we introduce a simple and effective predictive reconstruction objective that unifies predictive learning with video reconstruction. Specifically, we randomly discard future frames and encode only partial past observations, while training the decoder to reconstruct the observed frames and predict future ones simultaneously. This design encourages the latent space to encode temporally predictive structures and build a more coherent understanding of video dynamics, thereby improving generation quality. Our model, termed Predictive Video VAE (PV-VAE), achieves superior performance on video generation, with 52% faster convergence and a 34.42 FVD improvement over the Wan2.2 VAE on UCF101. Furthermore, comprehensive analyses demonstrate that PV-VAE not only exhibits favorable scalability, with generative performance improving alongside VAE training, but also yields consistent gains in downstream video understanding, underscoring a latent space that effectively captures temporal coherence and motion priors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。