arXiv:2507.23785cs.CV2025-07ICCV被引 32

用视频生成高保真动态3D内容,无需逐个建模。

Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis

  • 通过高斯点场变分自编码器压缩3D动态数据到紧凑潜空间。
  • 在Objaverse数据集上训练,生成质量优于现有方法。
  • 仅用合成数据训练,却能处理真实视频,泛化能力强。

本文提出一种新颖的视频到4D生成框架,可从单个视频输入生成高质量动态3D内容。直接进行4D扩散建模因数据构建成本高且表示维度庞大而极具挑战。我们引入一种直接的4DMesh-to-GS变分场变分自编码器(VAE),无需实例级拟合即可从3D动画数据中编码规范高斯点(GS)及其时序变化,并将高维动画压缩至紧凑潜空间。在此高效表征基础上,我们训练了一个基于时序感知扩散变换器(Temporal-aware Diffusion Transformer)的高斯变分场扩散模型,以输入视频和规范高斯点为条件。模型在精心筛选的Objaverse数据集上的可动画3D物体上训练,生成质量显著优于现有方法。即使仅在合成数据上训练,仍表现出对真实场景视频的强大泛化能力,为高质量动画3D内容生成开辟新路径。

原文摘要 · Abstract (English)

In this paper, we present a novel framework for video-to-4D generation that creates high-quality dynamic 3D content from single video inputs. Direct 4D diffusion modeling is extremely challenging due to costly data construction and the high-dimensional nature of jointly representing 3D shape, appearance, and motion. We address these challenges by introducing a Direct 4DMesh-to-GS Variation Field VAE that directly encodes canonical Gaussian Splats (GS) and their temporal variations from 3D animation data without per-instance fitting, and compresses high-dimensional animations into a compact latent space. Building upon this efficient representation, we train a Gaussian Variation Field diffusion model with temporal-aware Diffusion Transformer conditioned on input videos and canonical GS. Trained on carefully-curated animatable 3D objects from the Objaverse dataset, our model demonstrates superior generation quality compared to existing methods. It also exhibits remarkable generalization to in-the-wild video inputs despite being trained exclusively on synthetic data, paving the way for generating high-quality animated 3D content. Project page: https://gvfdiffusion.github.io/.

视频生成4D重建扩散模型高斯点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。