arXiv:2508.13154cs.CV2025-08被引 34

单图生成动态3D场景,效率远超传统方法。

4DNeX: Feed-Forward 4D Generative Modeling Made Easy

  • 用预训练视频扩散模型微调,实现端到端单图转4D。
  • 在4DNeX-10M数据集上生成高质量动态点云,支持新视角视频合成。
  • 适合做3D动态生成、虚拟世界建模的研究与应用开发者。

我们提出4DNeX,首个基于前馈架构的单图生成4D(即动态3D)场景表示框架。不同于依赖复杂优化或多帧视频输入的现有方法,4DNeX通过微调预训练视频扩散模型,实现高效端到端图像到4D生成。首先,为缓解4D数据稀缺,构建了包含1000万条高质量4D标注的4DNeX-10M数据集,采用先进重建方法生成。其次,提出统一的6D视频表示,联合建模RGB与XYZ序列,促进外观与几何结构化学习。第三,设计一系列简单有效的适配策略,将预训练视频扩散模型成功迁移到4D建模任务。4DNeX生成高质量动态点云,支持新视角视频合成。大量实验表明,其在效率和泛化能力上优于现有4D生成方法,为单图到4D建模提供可扩展解决方案,并为生成式4D世界模型模拟动态场景演化奠定基础。

原文摘要 · Abstract (English)

We present 4DNeX, the first feed-forward framework for generating 4D (i.e., dynamic 3D) scene representations from a single image. In contrast to existing methods that rely on computationally intensive optimization or require multi-frame video inputs, 4DNeX enables efficient, end-to-end image-to-4D generation by fine-tuning a pretrained video diffusion model. Specifically, 1) to alleviate the scarcity of 4D data, we construct 4DNeX-10M, a large-scale dataset with high-quality 4D annotations generated using advanced reconstruction approaches. 2) we introduce a unified 6D video representation that jointly models RGB and XYZ sequences, facilitating structured learning of both appearance and geometry. 3) we propose a set of simple yet effective adaptation strategies to repurpose pretrained video diffusion models for 4D modeling. 4DNeX produces high-quality dynamic point clouds that enable novel-view video synthesis. Extensive experiments demonstrate that 4DNeX outperforms existing 4D generation methods in efficiency and generalizability, offering a scalable solution for image-to-4D modeling and laying the foundation for generative 4D world models that simulate dynamic scene evolution.

4D生成动态点云视频扩散单图生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。