arXiv:2412.04471cs.CVcs.AI2024-12被引 3

用文本生成逼真动态4D场景,支持任意视角观看。

PaintScene4D: Consistent 4D Scene Generation from Text Prompts

  • 基于真实视频数据训练的模型,不依赖3D合成数据
  • 通过渐进式扭曲与修复保持多视角时空一致性
  • 无需训练即可生成可自由调控相机轨迹的4D场景

扩散模型推动了2D和3D内容生成的发展,但生成逼真动态4D场景仍是难题。现有方法多依赖在合成物体数据集上微调的预训练3D生成模型,导致场景以物体为中心且缺乏真实感。而文本到视频模型虽能生成更真实动态,却在空间理解与相机视角控制方面表现不足。为此,我们提出PaintScene4D,一种全新的文本到4D场景生成框架。该框架采用轻量级架构,利用在多样化真实世界数据集上训练的视频生成模型。首先通过视频生成模型创建参考视频,再通过策略性相机阵列选择进行渲染。采用渐进式扭曲与修补技术,确保多视角间的空间与时间一致性。最后通过动态渲染器优化多视图图像,实现基于用户偏好的灵活相机控制。整个框架无需训练,可高效生成可从任意轨迹观看的逼真4D场景。代码将公开。项目页:https://paintscene4d.github.io/

原文摘要 · Abstract (English)

Recent advances in diffusion models have revolutionized 2D and 3D content creation, yet generating photorealistic dynamic 4D scenes remains a significant challenge. Existing dynamic 4D generation methods typically rely on distilling knowledge from pre-trained 3D generative models, often fine-tuned on synthetic object datasets. Consequently, the resulting scenes tend to be object-centric and lack photorealism. While text-to-video models can generate more realistic scenes with motion, they often struggle with spatial understanding and provide limited control over camera viewpoints during rendering. To address these limitations, we present PaintScene4D, a novel text-to-4D scene generation framework that departs from conventional multi-view generative models in favor of a streamlined architecture that harnesses video generative models trained on diverse real-world datasets. Our method first generates a reference video using a video generation model, and then employs a strategic camera array selection for rendering. We apply a progressive warping and inpainting technique to ensure both spatial and temporal consistency across multiple viewpoints. Finally, we optimize multi-view images using a dynamic renderer, enabling flexible camera control based on user preferences. Adopting a training-free architecture, our PaintScene4D efficiently produces realistic 4D scenes that can be viewed from arbitrary trajectories. The code will be made publicly available. Our project page is at https://paintscene4d.github.io/

4D生成文本生成视频生成相机控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。