用冻结视觉模型构建更优视频生成潜空间,提升语义组织与生成效率。
V-RAE: Rethinking Video Latent Spaces for Generation

- 基于冻结视觉模型构建轻量时序压缩潜空间,保留语义结构。
- 在K600上达2.13 rFVD,生成收敛速度提升6倍,且语义信息显著增强。
- 提出tFVD诊断指标,更好反映生成质量,适合追求高效生成的研究者。
视频生成依赖自编码器定义紧凑的潜空间供生成模型使用。尽管视频自编码器架构已大幅演进,其潜空间仍主要优化于像素级重建,缺乏高层次语义组织。然而,重建最优的潜空间未必适用于生成建模。我们提出V-RAE,一种基于冻结视觉基础模型表示构建紧凑生成潜空间的视频表示自编码器。轻量级时序池化模块消除时间冗余并保留语义结构,视频解码器从压缩特征重建连续运动。我们在四个代表性冻结编码器上评估V-RAE在视频重建、语义探测和类别条件生成上的表现。V-RAE在K600上达到2.13 rFVD,优于所有对比的大规模预训练视频VAE。其潜空间保留的语义信息远超传统视频分词器。在相同生成设置下,最优变体在UCF101和K600上分别取得117.86和19.16 gFVD,收敛速度最快提升6倍。我们进一步表明,重建质量不足以衡量生成效用,提出tFVD(时间一致性诊断),与下游生成质量相关性更强。此外,V-RAE在城市景观未来视频预测任务中也优于Wan 2.2 VAE潜空间。结果表明,冻结语义表示可支持视频重建、生成与预测建模。
原文摘要 · Abstract (English)
Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization. A reconstruction-optimal latent space, however, need not be well suited to generative modeling. We propose V-RAE, a video representation autoencoder that builds compact generative latents on top of frozen vision foundation model representations. A lightweight temporal pooling module removes temporal redundancy while preserving semantic structure, and a video decoder reconstructs continuous motion from the compressed features. We evaluate V-RAE with four representative frozen encoders on video reconstruction, semantic probing, and class-conditional generation. V-RAE achieves 2.13 rFVD on K600, outperforming all evaluated large-scale pretrained video VAEs. Its latents retain substantially more semantic information than conventional video tokenizer latents. Under matched generation settings, our best variant achieves gFVD scores of 117.86 and 19.16 on UCF101 and K600, respectively, while converging up to 6x faster}. We further show that reconstruction quality alone is insufficient to characterize generative utility and introduce tFVD, a temporal-coherence diagnostic that correlates more reliably with downstream generation quality. Beyond video generation, V-RAE also improves future video prediction on Cityscapes over the Wan 2.2 VAE latent space under matched prediction settings. Taken together, the experiments show that frozen semantic representations can support video reconstruction, generation, and predictive modeling. The project page: https://v-rae.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。