arXiv:2411.18613cs.CV2024-11CVPR被引 175

用单视频生成可任意视角观看的动态3D场景

CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models

论文配图:CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models
图 1 · 摘自论文原文
  • 基于多视角视频扩散模型,从单视频生成多视角内容
  • 支持任意相机位姿和时间点的新视角合成,重建效果优异
  • 适合影视创作与虚拟场景生成,交互式演示可在线体验

我们提出CAT4D,一种从单视角视频生成4D(动态3D)场景的方法。CAT4D利用在多样化数据集组合上训练的多视角视频扩散模型,可在任意指定相机位姿和时间点实现新视角合成。结合一种新型采样策略,该模型可将单个单视角视频转换为多视角视频,通过优化可变形3D高斯表示实现鲁棒的4D重建。我们在新视角合成与动态场景重建基准上展示了具有竞争力的表现,并凸显了从真实或生成视频生成4D场景的创造能力。更多结果与交互演示请见项目主页:https://cat-4d.github.io/。

原文摘要 · Abstract (English)

We present CAT4D, a method for creating 4D (dynamic 3D) scenes from monocular video. CAT4D leverages a multi-view video diffusion model trained on a diverse combination of datasets to enable novel view synthesis at any specified camera poses and timestamps. Combined with a novel sampling approach, this model can transform a single monocular video into a multi-view video, enabling robust 4D reconstruction via optimization of a deformable 3D Gaussian representation. We demonstrate competitive performance on novel view synthesis and dynamic scene reconstruction benchmarks, and highlight the creative capabilities for 4D scene generation from real or generated videos. See our project page for results and interactive demos: https://cat-4d.github.io/.

4D生成视频扩散动态重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。