arXiv:2603.22972cs.CV2026-03中稿 · ECCV被引 2

用网格骨架控制图像生成,实现可导航的多房间3D场景。

WorldMesh: Generating Navigable Multi-Room 3D Scenes via Mesh-Conditioned Image Diffusion

论文配图:WorldMesh: Generating Navigable Multi-Room 3D Scenes via Mesh-Conditioned Image Diffusion
图 1 · 摘自论文原文
  • 先生成场景几何网格,再用图像模型基于网格生成真实外观。
  • 支持任意大小场景,物体丰富且布局合理,保持3D一致性。
  • 适合需要高保真3D环境的应用,如虚拟现实和游戏开发。

近期图像与视频合成进展推动了3D场景生成的发展。然而,我们发现文本到图像/视频的方法在缺乏持续显式几何表示时,难以在较大环境尺度上维持场景与物体的一致性。为此,我们提出一种以几何为核心的方案,将大规模3D场景生成分解为结构构建(以网格骨架表示)和外观合成两部分。给定文本描述后,首先构建包含墙壁、地板等元素的几何网格,再通过图像合成、分割与物体重建,在网格中填充符合真实布局的物体。该网格骨架被渲染为条件输入,驱动图像合成,提供一致的外观生成结构支撑。该方法可生成任意规模、物体丰富多样、具备强3D一致性的高质量3D场景,显著推进了环境级沉浸式3D世界生成的实现。

原文摘要 · Abstract (English)

Recent progress in image and video synthesis has inspired their use in advancing 3D scene generation. However, we observe that text-to-image and -video approaches struggle to maintain scene- and object-level consistency beyond a limited environment scale without a persistent, explicit geometric representation. We thus present a geometry-first approach that decouples this complex problem of large-scale 3D scene synthesis into its structural composition, represented as a mesh scaffold, and realistic appearance synthesis, which leverages powerful image synthesis models conditioned on the mesh scaffold. From an input text description, we first construct a mesh capturing the environment's geometry (walls, floors, etc.), and then use image synthesis, segmentation and object reconstruction to populate the mesh structure with objects in realistic layouts. This mesh scaffold is then rendered to condition image synthesis, providing a structural backbone for consistent appearance generation. This enables scalable, arbitrarily-sized 3D scenes of high object richness and diversity, combining robust 3D consistency with photorealistic detail. We believe this marks a significant step toward generating truly environment-scale, immersive 3D worlds.

3D生成图像扩散网格建模场景合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。