单图生成动态3D场景,效率远超传统方法。
4DNeX: Feed-Forward 4D Generative Modeling Made Easy
- 用预训练视频扩散模型微调,实现端到端单图转4D。
- 在4DNeX-10M数据集上生成高质量动态点云,支持新视角视频合成。
- 适合做3D动态生成、虚拟世界建模的研究与应用开发者。
我们提出4DNeX,首个基于前馈架构的单图生成4D(即动态3D)场景表示框架。不同于依赖复杂优化或多帧视频输入的现有方法,4DNeX通过微调预训练视频扩散模型,实现高效端到端图像到4D生成。首先,为缓解4D数据稀缺,构建了包含1000万条高质量4D标注的4DNeX-10M数据集,采用先进重建方法生成。其次,提出统一的6D视频表示,联合建模RGB与XYZ序列,促进外观与几何结构化学习。第三,设计一系列简单有效的适配策略,将预训练视频扩散模型成功迁移到4D建模任务。4DNeX生成高质量动态点云,支持新视角视频合成。大量实验表明,其在效率和泛化能力上优于现有4D生成方法,为单图到4D建模提供可扩展解决方案,并为生成式4D世界模型模拟动态场景演化奠定基础。
原文摘要 · Abstract (English)
We present 4DNeX, the first feed-forward framework for generating 4D (i.e., dynamic 3D) scene representations from a single image. In contrast to existing methods that rely on computationally intensive optimization or require multi-frame video inputs, 4DNeX enables efficient, end-to-end image-to-4D generation by fine-tuning a pretrained video diffusion model. Specifically, 1) to alleviate the scarcity of 4D data, we construct 4DNeX-10M, a large-scale dataset with high-quality 4D annotations generated using advanced reconstruction approaches. 2) we introduce a unified 6D video representation that jointly models RGB and XYZ sequences, facilitating structured learning of both appearance and geometry. 3) we propose a set of simple yet effective adaptation strategies to repurpose pretrained video diffusion models for 4D modeling. 4DNeX produces high-quality dynamic point clouds that enable novel-view video synthesis. Extensive experiments demonstrate that 4DNeX outperforms existing 4D generation methods in efficiency and generalizability, offering a scalable solution for image-to-4D modeling and laying the foundation for generative 4D world models that simulate dynamic scene evolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。