arXiv:2510.23880cs.CVcs.GR2025-10被引 8

用文字生成的物体模型拼出360度可看的大场景,无需训练

TRELLISWorld: Training-Free World Generation from Object Generators

  • 把文本生成3D物体模型当积木,分块生成再融合成大场景
  • 无需场景数据或重训练,能生成多样布局且局部语义可控
  • 适合快速构建虚拟场景,尤其适合没专业3D技能的用户

文本驱动的3D场景生成在虚拟原型设计、AR/VR和仿真中具有广阔应用前景。然而,现有方法常受限于单物体生成、需特定领域训练或缺乏360度视角支持。本文提出一种无需训练的3D场景合成方法,通过复用通用文本到3D物体扩散模型作为模块化图块生成器。我们将场景生成重构为多图块去噪问题,独立生成重叠的3D区域,并通过加权平均实现无缝融合。该方法支持大规模、连贯场景的可扩展生成,同时保持局部语义控制。本方法无需场景级数据集或重训练,仅依赖少量启发式规则,继承了物体级先验的泛化能力。实验表明,该方法支持多样化场景布局、高效生成与灵活编辑,为通用语言驱动3D场景构建提供了简单而强大的基础。

原文摘要 · Abstract (English)

Text-driven 3D scene generation holds promise for a wide range of applications, from virtual prototyping to AR/VR and simulation. However, existing methods are often constrained to single-object generation, require domain-specific training, or lack support for full 360-degree viewability. In this work, we present a training-free approach to 3D scene synthesis by repurposing general-purpose text-to-3D object diffusion models as modular tile generators. We reformulate scene generation as a multi-tile denoising problem, where overlapping 3D regions are independently generated and seamlessly blended via weighted averaging. This enables scalable synthesis of large, coherent scenes while preserving local semantic control. Our method eliminates the need for scene-level datasets or retraining, relies on minimal heuristics, and inherits the generalization capabilities of object-level priors. We demonstrate that our approach supports diverse scene layouts, efficient generation, and flexible editing, establishing a simple yet powerful foundation for general-purpose, language-driven 3D scene construction.

3D生成文本生成场景构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。