arXiv:2501.01722cs.CV2025-01被引 19

无需SDS,通过自回归方式生成更连贯的4D动态3D内容。

AR4D: Autoregressive 4D Generation from Monocular Videos

论文配图:AR4D: Autoregressive 4D Generation from Monocular Videos
图 1 · 摘自论文原文
  • 采用自回归框架,逐帧生成3D表示以提升运动与几何准确性。
  • 在无SDS条件下实现更优多样性、时空一致性和提示对齐效果。
  • 适合需要高质量动态3D生成的研究者或内容创作者。

近期生成模型的发展推动了动态3D内容生成(即4D生成)的热潮。现有方法多依赖得分蒸馏采样(SDS)推断新视角视频,常因SDS固有的随机性导致多样性有限、时空不一致及提示对齐不佳。为此,我们提出AR4D,一种无需SDS的新型4D生成范式。该范式包含三个阶段:首先,利用预训练专家模型对单目视频首帧生成3D表示,并微调为规范空间;其次,借鉴视频自然的自回归特性,基于前一帧生成当前帧的3D表示,以提升几何与运动估计精度;同时引入渐进式视图采样策略,利用大规模预训练3D重建模型的先验防止过拟合;最后,为避免自回归生成带来的外观漂移,引入基于全局变形场和每帧3D几何的精修阶段。大量实验表明,AR4D在无需SDS的情况下达到当前最优4D生成性能,显著提升多样性、时空一致性与提示对齐能力。

原文摘要 · Abstract (English)

Recent advancements in generative models have ignited substantial interest in dynamic 3D content creation (\ie, 4D generation). Existing approaches primarily rely on Score Distillation Sampling (SDS) to infer novel-view videos, typically leading to issues such as limited diversity, spatial-temporal inconsistency and poor prompt alignment, due to the inherent randomness of SDS. To tackle these problems, we propose AR4D, a novel paradigm for SDS-free 4D generation. Specifically, our paradigm consists of three stages. To begin with, for a monocular video that is either generated or captured, we first utilize pre-trained expert models to create a 3D representation of the first frame, which is further fine-tuned to serve as the canonical space. Subsequently, motivated by the fact that videos happen naturally in an autoregressive manner, we propose to generate each frame's 3D representation based on its previous frame's representation, as this autoregressive generation manner can facilitate more accurate geometry and motion estimation. Meanwhile, to prevent overfitting during this process, we introduce a progressive view sampling strategy, utilizing priors from pre-trained large-scale 3D reconstruction models. To avoid appearance drift introduced by autoregressive generation, we further incorporate a refinement stage based on a global deformation field and the geometry of each frame's 3D representation. Extensive experiments have demonstrated that AR4D can achieve state-of-the-art 4D generation without SDS, delivering greater diversity, improved spatial-temporal consistency, and better alignment with input prompts.

4D生成自回归3D重建单目视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。