从单图生成完整3D场景,自动优化结构与纹理。
Self-Evolving 3D Scene Generation from a Single Image

- 融合3D生成与视频模型优势,分三阶段迭代优化
- 在复杂场景中实现几何稳定、视图一致纹理
- 无需训练,适合需快速生成3D资产的开发者
从单张图像生成高质量、带纹理的3D场景仍是视觉与图形学中的核心挑战。现有图像到3D生成模型虽能从单视角恢复合理几何结构,但其以物体为中心的训练方式限制了对复杂大尺度场景的泛化能力。我们提出EvoScene,一种自演化、免训练的框架,可逐步重构完整的3D场景。核心思想是结合现有模型的优势:利用3D生成模型进行几何推理,借助视频生成模型获取视觉知识。通过三个迭代阶段——空间先验初始化、视觉引导的3D场景网格生成、空间引导的新视角生成——EvoScene在2D与3D域间交替优化,逐步提升结构与外观质量。在多样化场景上的实验表明,EvoScene在几何稳定性、视图一致性纹理和未见区域补全方面均优于强基线,生成可用于实际应用的3D网格。
原文摘要 · Abstract (English)
Generating high-quality, textured 3D scenes from a single image remains a fundamental challenge in vision and graphics. Recent image-to-3D generators recover reasonable geometry from single views, but their object-centric training limits generalization to complex, large-scale scenes with faithful structure and texture. We present EvoScene, a self-evolving, training-free framework that progressively reconstructs complete 3D scenes from single images. The key idea is combining the complementary strengths of existing models: geometric reasoning from 3D generation models and visual knowledge from video generation models. Through three iterative stages--Spatial Prior Initialization, Visual-guided 3D Scene Mesh Generation, and Spatial-guided Novel View Generation--EvoScene alternates between 2D and 3D domains, gradually improving both structure and appearance. Experiments on diverse scenes demonstrate that EvoScene achieves superior geometric stability, view-consistent textures, and unseen-region completion compared to strong baselines, producing ready-to-use 3D meshes for practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。