arXiv:2411.04928cs.CVcs.AI2024-11被引 151

单图生成逼真3D/4D场景,可控视频扩散技术突破

DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion

  • 用时空解耦的视频扩散模型,分离空间与时间控制
  • 在Real-World、Synthetic数据集上超越现有方法,支持精细调控
  • 适合影视特效、虚拟现实开发者,实现单图生成动态3D场景

本文提出DimensionX,一种仅需单张图像即可生成逼真3D与4D场景的框架。核心思路是将3D空间结构与4D时间演化通过视频帧序列建模。尽管近期视频扩散模型生成效果生动,但因生成过程缺乏空间与时间可控性,难以直接还原3D/4D场景。为此,我们提出ST-Director,通过从维度差异数据中学习维度感知的LoRA,解耦视频扩散中的空间与时间因素。该可控生成机制使我们能精准操控空间结构与时间动态,结合多维信息重建3D与4D表示。此外,为提升生成质量,引入轨迹感知机制增强3D生成,采用身份保持去噪策略优化4D生成。在多个真实世界与合成数据集上的大量实验表明,DimensionX在可控视频生成、3D与4D场景生成方面均优于先前方法。

原文摘要 · Abstract (English)

In this paper, we introduce \textbf{DimensionX}, a framework designed to generate photorealistic 3D and 4D scenes from just a single image with video diffusion. Our approach begins with the insight that both the spatial structure of a 3D scene and the temporal evolution of a 4D scene can be effectively represented through sequences of video frames. While recent video diffusion models have shown remarkable success in producing vivid visuals, they face limitations in directly recovering 3D/4D scenes due to limited spatial and temporal controllability during generation. To overcome this, we propose ST-Director, which decouples spatial and temporal factors in video diffusion by learning dimension-aware LoRAs from dimension-variant data. This controllable video diffusion approach enables precise manipulation of spatial structure and temporal dynamics, allowing us to reconstruct both 3D and 4D representations from sequential frames with the combination of spatial and temporal dimensions. Additionally, to bridge the gap between generated videos and real-world scenes, we introduce a trajectory-aware mechanism for 3D generation and an identity-preserving denoising strategy for 4D generation. Extensive experiments on various real-world and synthetic datasets demonstrate that DimensionX achieves superior results in controllable video generation, as well as in 3D and 4D scene generation, compared with previous methods.

3D生成视频扩散可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。