arXiv:2503.20785cs.CV2025-03ICCV被引 40

单图生成4D场景,无需微调且时空一致

Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency

论文配图:Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency
图 1 · 摘自论文原文
  • 用预训练模型蒸馏生成初始4D结构,避免大规模数据训练
  • 通过点引导去噪和潜在变量替换,实现画面间时空一致性
  • 可实时可控渲染,适合影视动画等需要快速生成的场景

我们提出Free4D,一种从单张图像出发、无需微调的4D场景生成框架。现有方法或局限于物体级生成,无法实现场景级建模,或依赖大规模多视角视频数据进行昂贵训练,且因4D数据稀缺导致泛化能力差。我们的核心思路是利用预训练基础模型蒸馏出一致的4D场景表示,具有高效性与强泛化优势。首先,使用图像到视频扩散模型动画化输入图像,并初始化4D几何结构;其次,设计自适应引导机制:通过点引导去噪策略保证空间一致性,提出新颖的潜在变量替换策略提升时间连贯性;最后,提出基于调制的精修方法,在充分保留生成信息的同时缓解不一致问题。最终生成的4D表示支持实时、可控渲染,显著推进了单图驱动的4D场景生成技术。

原文摘要 · Abstract (English)

We present Free4D, a novel tuning-free framework for 4D scene generation from a single image. Existing methods either focus on object-level generation, making scene-level generation infeasible, or rely on large-scale multi-view video datasets for expensive training, with limited generalization ability due to the scarcity of 4D scene data. In contrast, our key insight is to distill pre-trained foundation models for consistent 4D scene representation, which offers promising advantages such as efficiency and generalizability. 1) To achieve this, we first animate the input image using image-to-video diffusion models followed by 4D geometric structure initialization. 2) To turn this coarse structure into spatial-temporal consistent multiview videos, we design an adaptive guidance mechanism with a point-guided denoising strategy for spatial consistency and a novel latent replacement strategy for temporal coherence. 3) To lift these generated observations into consistent 4D representation, we propose a modulation-based refinement to mitigate inconsistencies while fully leveraging the generated information. The resulting 4D representation enables real-time, controllable rendering, marking a significant advancement in single-image-based 4D scene generation.

4D生成扩散模型单图生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。