arXiv:2605.12957cs.CV2026-05IJCV被引 1

先建几何再加细节,让单图生成的3D世界更真实一致。

GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion

论文配图:GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video Diffusion
图 1 · 摘自论文原文
  • 分两阶段生成:先用视频扩散模型建粗略几何结构,再根据几何补全精细外观。
  • 在多个数据集上生成的3D场景几何准确度和跨视角一致性显著优于现有方法。
  • 适合需要高质量3D内容的自动驾驶、虚拟现实等应用,且训练数据需求少。

生成模型与大规模数据集的发展推动了3D世界生成的进步,广泛应用于空间智能、具身智能和自动驾驶等领域。然而,现有方法多侧重外观预测,对底层几何建模不足,导致场景结构不可靠、跨视角一致性差。受人类视觉从粗到细的感知机制启发,我们提出GTA——一种遵循‘几何先于外观’范式的图像到3D世界生成新方法。给定单张输入图像,GTA采用双阶段框架,使用两个专用视频扩散模型:第一阶段从新视角生成粗略几何结构;第二阶段基于预测几何合成精细外观。为增强跨视角外观一致性,训练时引入随机潜变量打乱策略,并在测试时采用缩放方案提升感知质量而不牺牲定量指标。大量实验表明,GTA在保真度、视觉质量和几何准确性方面均显著优于现有方法。此外,GTA可作为通用增强模块,提升已有图像到3D生成流程的质量,支持多种下游应用,且训练时表现出良好数据效率,体现其通用性与广泛适用性。

原文摘要 · Abstract (English)

Recent developments in generative models and large-scale datasets have substantially advanced 3D world generation, facilitating a broad range of domains including spatial intelligence, embodied intelligence, and autonomous driving. While achieving remarkable progress, existing approaches to 3D world generation typically prioritize appearance prediction with limited modeling of the underlying geometry, leading to issues such as unreliable scene structure estimation and degraded cross-view consistency. To address these limitations, motivated by the coarse-to-fine nature of human visual perception, we propose GTA, a novel image-to-3D world generation method following a Geometry-Then-Appearance paradigm. Specifically, given a single input image, to improve the structural fidelity of synthesized 3D scenes, GTA adopts a two-stage framework with two dedicated video diffusion models, which first generate coarse geometric structure from novel viewpoints and then synthesize fine-grained appearance conditioned on the predicted geometry. To further enhance cross-view appearance consistency, we introduce a random latent shuffle strategy during the training process, along with a test-time scaling scheme that improves perceptual quality without compromising quantitative performance. Extensive experiments have demonstrated that our proposed method consistently outperforms existing approaches in terms of fidelity, visual quality, and geometric accuracy. Moreover, GTA is shown to be effective as a general enhancement module that further improves the generation quality of existing image-to-3D world pipelines, as well as supporting multiple downstream applications and exhibiting favorable data efficiency during model training, highlighting its versatility and broad applicability. Project page: https://hanxinzhu-lab.github.io/GTA/.

3D生成视频扩散几何建模图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。