文本生成复杂动态4D场景,区分前景运动与背景变化。
CoCo4D: Comprehensive and Complex 4D Scene Generation
- 分两阶段生成:先建模动态前景,再构建变化背景。
- 通过渐进式外扩实现多视角一致的4D场景合成。
- 支持文本或图文输入,适合影视、游戏场景生成。
现有4D合成方法主要聚焦物体级生成或有限新视角的动态场景合成,难以生成多视角一致且沉浸感强的动态4D场景。为此,我们提出一种名为CoCo4D的框架,可基于文本提示(支持附加图像)生成细节丰富的动态4D场景。该方法观察到:关节运动通常表征前景物体,而背景变化较弱。因此,CoCo4D将4D场景合成分为两个任务:建模动态前景和生成演化背景,均由参考运动序列引导。给定文本提示和可选参考图像后,首先利用视频扩散模型生成初始运动序列,再通过新颖的渐进式外扩方案合成前景与背景。为确保前景在动态背景中自然融合,CoCo4D优化了前景的参数化轨迹,实现真实且连贯的融合。大量实验表明,CoCo4D在4D场景生成上达到或优于现有方法,验证了其有效性与高效性。更多结果见官网 https://colezwhy.github.io/coco4d/。
原文摘要 · Abstract (English)
Existing 4D synthesis methods primarily focus on object-level generation or dynamic scene synthesis with limited novel views, restricting their ability to generate multi-view consistent and immersive dynamic 4D scenes. To address these constraints, we propose a framework (dubbed as CoCo4D) for generating detailed dynamic 4D scenes from text prompts, with the option to include images. Our method leverages the crucial observation that articulated motion typically characterizes foreground objects, whereas background alterations are less pronounced. Consequently, CoCo4D divides 4D scene synthesis into two responsibilities: modeling the dynamic foreground and creating the evolving background, both directed by a reference motion sequence. Given a text prompt and an optional reference image, CoCo4D first generates an initial motion sequence utilizing video diffusion models. This motion sequence then guides the synthesis of both the dynamic foreground object and the background using a novel progressive outpainting scheme. To ensure seamless integration of the moving foreground object within the dynamic background, CoCo4D optimizes a parametric trajectory for the foreground, resulting in realistic and coherent blending. Extensive experiments show that CoCo4D achieves comparable or superior performance in 4D scene generation compared to existing methods, demonstrating its effectiveness and efficiency. More results are presented on our website https://colezwhy.github.io/coco4d/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。