arXiv:2412.16253cs.CVcs.GR2024-12

仅用一段实拍视频,10分钟内生成可控的逼真3D场景。

ExCellGen: Fast, Controllable, Photorealistic 3D Scene Generation from a Single Real-World Exemplar

  • 输入单段实拍视频,用3D高斯溅射重建场景外观。
  • 训练局部细胞自动机生成稀疏体素,实现0.5-2秒快速生成。
  • 支持交互式编辑,适合影视建模与数字孪生开发者使用。

逼真3D场景生成因缺乏大规模高质量真实世界3D数据集及复杂手工建模流程而困难重重,导致迭代缓慢,限制创造力。本文提出ExCellGen框架,仅需一段手持视频或无人机影像即可快速生成3D场景。首先利用3D高斯溅射(3DGS)对输入场景进行鲁棒重建,获得高质量3D外观模型;随后训练每场景专用的生成式细胞自动机(GCA),生成特征化稀疏体素,有效分摊生成成本并支持可控性;最后通过基于补丁的重映射步骤,从原始3D高斯溅射中合成完整场景,成功恢复输入场景的外观统计特性。整个流程每个示例训练时间小于10分钟,生成时间仅0.5-2秒。系统支持全用户交互控制,并在自包含交互式GUI中展示了从真实世界示例生成复杂3D场景的效果。

原文摘要 · Abstract (English)

Photorealistic 3D scene generation is challenging due to the scarcity of large-scale, high-quality real-world 3D datasets and complex workflows requiring specialized expertise for manual modeling. These constraints often result in slow iteration cycles, where each modification demands substantial effort, ultimately stifling creativity. We propose a fast, exemplar-driven framework for generating 3D scenes from a single casual input, such as handheld video or drone footage. Our method first leverages 3D Gaussian Splatting (3DGS) to robustly reconstruct input scenes with a high-quality 3D appearance model. We then train a per-scene Generative Cellular Automaton (GCA) to produce a sparse volume of featurized voxels, effectively amortizing scene generation while enabling controllability. A subsequent patch-based remapping step composites the complete scene from the exemplar's initial 3D Gaussian splats, successfully recovering the appearance statistics of the input scene. The entire pipeline can be trained in less than 10 minutes for each exemplar and generates scenes in 0.5-2 seconds. Our method enables interactive creation with full user control, and we showcase complex 3D generation results from real-world exemplars within a self-contained interactive GUI.

3D生成实时渲染视觉重建交互设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。