arXiv:2411.02319cs.CVcs.AI2024-11ICLR被引 54

提出可生成任意3D与4D场景的GenXD框架,突破数据与模型瓶颈。

GenXD: Generating Any 3D and 4D Scenes

  • 利用日常视频中的相机与物体运动,构建多视角时序模块解耦运动
  • 构建首个大规模真实世界4D数据集CamVid-30K,支持3D/4D联合生成
  • 支持轨迹跟随视频与可转为3D表示的一致视图,适合场景生成研究者

2D视觉生成已取得显著进展,但3D与4D生成在真实应用中仍具挑战,主要受限于缺乏大规模4D数据和有效模型设计。本文通过利用日常生活中常见的相机与物体运动,联合研究通用3D与4D生成。由于社区缺乏真实世界4D数据,我们首先提出一种数据整理流程,从视频中提取相机位姿与物体运动强度。基于此流程,构建大规模真实世界4D场景数据集CamVid-30K。结合所有3D与4D数据,我们开发了GenXD框架,可生成任意3D或4D场景。提出多视角时序模块,解耦相机与物体运动,实现对3D与4D数据的无缝学习。此外,GenXD采用掩码潜在条件,支持多种条件视图。实验表明,其可生成沿相机轨迹的视频以及可提升为3D表示的一致3D视图,在多个真实与合成数据集上优于先前方法。

原文摘要 · Abstract (English)

Recent developments in 2D visual generation have been remarkably successful. However, 3D and 4D generation remain challenging in real-world applications due to the lack of large-scale 4D data and effective model design. In this paper, we propose to jointly investigate general 3D and 4D generation by leveraging camera and object movements commonly observed in daily life. Due to the lack of real-world 4D data in the community, we first propose a data curation pipeline to obtain camera poses and object motion strength from videos. Based on this pipeline, we introduce a large-scale real-world 4D scene dataset: CamVid-30K. By leveraging all the 3D and 4D data, we develop our framework, GenXD, which allows us to produce any 3D or 4D scene. We propose multiview-temporal modules, which disentangle camera and object movements, to seamlessly learn from both 3D and 4D data. Additionally, GenXD employs masked latent conditions to support a variety of conditioning views. GenXD can generate videos that follow the camera trajectory as well as consistent 3D views that can be lifted into 3D representations. We perform extensive evaluations across various real-world and synthetic datasets, demonstrating GenXD's effectiveness and versatility compared to previous methods in 3D and 4D generation.

3D生成4D生成多视角场景建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。