用2D扩散模型生成纹理参考图,指导3D场景结构构建。
GeoDiff3D: Self-Supervised 3D Scene Generation with Geometry-Constrained 2D Diffusion Guidance
- 以粗略几何为锚点,用约束的2D扩散模型生成带纹理的参考图像。
- 在复杂场景中生成更高质量、细节更丰富的3D场景,减少对标注数据依赖。
- 适合游戏、影视和VR/AR领域快速生成高保真3D场景。
3D场景生成是游戏、影视/视觉特效及VR/AR的核心技术。随着对快速迭代、高保真细节和易用内容创作的需求增加,该领域受到更多关注。现有方法主要分为间接2D到3D重建和直接3D生成两类,但均受限于结构建模能力弱和对大规模真实标签数据的依赖,常导致结构伪影、几何不一致以及复杂场景中高频细节退化。我们提出GeoDiff3D,一种高效的自监督框架,利用粗略几何作为结构锚点,并通过几何约束的2D扩散模型提供富含纹理的参考图像。重要的是,GeoDiff3D不要求扩散生成参考图具备严格的多视角一致性,且对产生的噪声与不一致引导仍保持鲁棒性。我们进一步引入体素对齐的3D特征聚合与双重自监督机制,在显著降低对标注数据依赖的同时,维持场景连贯性和精细细节。GeoDiff3D训练成本低,支持快速生成高质量3D场景。在挑战性场景上的大量实验表明,其泛化能力和生成质量优于现有基线,为可访问、高效的3D场景构建提供了实用方案。
原文摘要 · Abstract (English)
3D scene generation is a core technology for gaming, film/VFX, and VR/AR. Growing demand for rapid iteration, high-fidelity detail, and accessible content creation has further increased interest in this area. Existing methods broadly follow two paradigms - indirect 2D-to-3D reconstruction and direct 3D generation - but both are limited by weak structural modeling and heavy reliance on large-scale ground-truth supervision, often producing structural artifacts, geometric inconsistencies, and degraded high-frequency details in complex scenes. We propose GeoDiff3D, an efficient self-supervised framework that uses coarse geometry as a structural anchor and a geometry-constrained 2D diffusion model to provide texture-rich reference images. Importantly, GeoDiff3D does not require strict multi-view consistency of the diffusion-generated references and remains robust to the resulting noisy, inconsistent guidance. We further introduce voxel-aligned 3D feature aggregation and dual self-supervision to maintain scene coherence and fine details while substantially reducing dependence on labeled data. GeoDiff3D also trains with low computational cost and enables fast, high-quality 3D scene generation. Extensive experiments on challenging scenes show improved generalization and generation quality over existing baselines, offering a practical solution for accessible and efficient 3D scene construction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。