仅用一张俯视图生成高质量3D场景,无需训练
Constructing a 3D Scene from a Single Image
- 分区域生成+空间感知修复,确保结构一致
- 在多个数据集上超越现有方法的几何质量和纹理保真度
- 无需3D标注或微调,适合快速构建3D资产
获取详细3D场景通常需要昂贵设备、多视角数据或人工建模。因此,从单张俯视图生成复杂3D场景的轻量级方法在实际应用中至关重要。尽管近期3D生成模型在物体级别表现优异,但扩展至完整场景时常出现几何不一致、布局幻觉和网格质量差的问题。本文提出SceneFuse-3D,一种无需训练的框架,可从单张俯视图合成连贯3D场景。方法基于两大原则:区域化生成以提升图像到3D的对齐与分辨率,以及空间感知3D补全以保证全局一致性与高质量几何生成。具体而言,将输入图像分解为重叠区域,使用预训练3D物体生成器分别生成各区域,再通过掩码修正流补全缺失几何,保持结构连续性。这种模块化设计克服了分辨率瓶颈,无需3D监督或微调即可保留空间结构。大量实验表明,SceneFuse-3D在几何质量、空间一致性和纹理保真度方面优于Trellis、Hunyuan3D-2、TripoSG和LGM等先进基线方法。结果证明,仅凭单张俯视图,通过原理性强、无需训练的流程即可实现高质量3D场景生成。
原文摘要 · Abstract (English)
Acquiring detailed 3D scenes typically demands costly equipment, multi-view data, or labor-intensive modeling. Therefore, a lightweight alternative, generating complex 3D scenes from a single top-down image, plays an essential role in real-world applications. While recent 3D generative models have achieved remarkable results at the object level, their extension to full-scene generation often leads to inconsistent geometry, layout hallucinations, and low-quality meshes. In this work, we introduce SceneFuse-3D, a training-free framework designed to synthesize coherent 3D scenes from a single top-down view. Our method is grounded in two principles: region-based generation to improve image-to-3D alignment and resolution, and spatial-aware 3D inpainting to ensure global scene coherence and high-quality geometry generation. Specifically, we decompose the input image into overlapping regions and generate each using a pretrained 3D object generator, followed by a masked rectified flow inpainting process that fills in missing geometry while maintaining structural continuity. This modular design allows us to overcome resolution bottlenecks and preserve spatial structure without requiring 3D supervision or fine-tuning. Extensive experiments across diverse scenes show that SceneFuse-3D outperforms state-of-the-art baselines, including Trellis, Hunyuan3D-2, TripoSG, and LGM, in terms of geometry quality, spatial coherence, and texture fidelity. Our results demonstrate that high-quality coherent 3D scene-level asset generation is achievable from a single top-down image using a principled, training-free pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。