arXiv:2608.10744cs.CV2026-08

用视频模型的潜在表示直接生成4D动态场景,无需重新训练。

Beyond Pixels: From Video Priors to 4D Worlds

论文配图:Beyond Pixels: From Video Priors to 4D Worlds
图 1 · 摘自论文原文
  • 利用共享VAE的视频潜在变量作为接口,直接映射到4D空间。
  • 在两个数据集上比现有方法提升2.88~5.81点,且人类评分更优。
  • 单个模型可跨多个视频生成器使用,无需调整或重训。

4D生成从文本或图像等条件中合成动态3D场景。现有方法要么用独立4D模型重建生成的RGB视频,要么将特定视频生成器改造成直接预测几何。前者存在分布不匹配与误差传播问题,后者则绑定特定生成器,更换生成器或条件方式需重新训练。本文提出:是否可利用共享变分自编码器(VAE)的视频模型最终去噪潜在变量,作为可复用的4D预测接口?基于此,我们提出直接潜向量到4D生成方法,即Latent-to-4D,跳过RGB中间步骤,通过对齐视频潜在变量与预训练4D解码器的标记网格,并结合帧内与全局时空注意力进行优化。仅在约1000个重建片段上训练,一个检查点即可在同VAE族内的多个视频扩散变换器间无修改迁移。在Text4D-200和I4D-200上,相比相同潜空间的Wan+4RC级联模型,投影式DINO-F1提升2.88–3.45和5.81点,同时在几何、时序稳定性与整体质量上更受人工评价青睐。

原文摘要 · Abstract (English)

4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a particular video generator to predict geometry directly. The former suffers from distribution mismatch and error propagation, whereas the latter ties 4D prediction to a specific generator and may require retraining when the generator or conditioning regime changes. We ask whether the final denoised latents of video models that share a variational autoencoder (VAE) can instead provide a reusable interface to explicit 4D prediction. Building on this insight, we introduce direct latent-to-4D generation and instantiate it as Latent-to-4D, which bypasses RGB by aligning a video latent with the token grid of a pretrained 4D decoder and refining it through frame-wise and global spatiotemporal attention. Trained on roughly 1K existing reconstruction clips, a single checkpoint transfers unchanged across multiple video diffusion transformers within the same VAE family. On Text4D-200 and I4D-200, Latent-to-4D surpasses matched same-latent Wan+4RC cascades in projection-based DINO-F1 by 2.88--3.45 and 5.81 points, respectively, while also being preferred by human raters for geometry, temporal stability, and overall quality.

4D生成视频生成扩散模型潜在空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。