arXiv:2510.15564cs.CV2025-10被引 14

用视觉引导生成高质量3D场景布局,解决内容贫乏与空间错乱问题。

Imaginarium: Vision-guided High-Quality 3D Scene Layout Generation

  • 通过图像生成模型扩展提示,对齐自建资产库
  • 基于视觉语义和几何信息解析3D布局,提升空间准确性
  • 结合场景图与整体视觉语义优化,保证逻辑一致性

生成艺术性强且连贯的3D场景布局在数字内容创作中至关重要。传统基于优化的方法受限于繁琐的手动规则,而深度生成模型难以产生丰富多样的内容。此外,使用大语言模型的方法常缺乏鲁棒性,无法准确捕捉复杂的空间关系。为此,本文提出一种新型视觉引导3D布局生成系统。首先构建包含2,037个场景资产和147个3D场景布局的高质量资产库。随后,采用图像生成模型将提示表示扩展为图像,并微调以匹配该资产库。接着开发鲁棒的图像解析模块,基于视觉语义与几何信息恢复场景的3D布局。最后,利用场景图与整体视觉语义优化布局,确保逻辑连贯性并匹配图像。大量用户测试表明,本方法在布局丰富度与质量上显著优于现有方法。代码与数据集将公开于https://github.com/HiHiAllen/Imaginarium。

原文摘要 · Abstract (English)

Generating artistic and coherent 3D scene layouts is crucial in digital content creation. Traditional optimization-based methods are often constrained by cumbersome manual rules, while deep generative models face challenges in producing content with richness and diversity. Furthermore, approaches that utilize large language models frequently lack robustness and fail to accurately capture complex spatial relationships. To address these challenges, this paper presents a novel vision-guided 3D layout generation system. We first construct a high-quality asset library containing 2,037 scene assets and 147 3D scene layouts. Subsequently, we employ an image generation model to expand prompt representations into images, fine-tuning it to align with our asset library. We then develop a robust image parsing module to recover the 3D layout of scenes based on visual semantics and geometric information. Finally, we optimize the scene layout using scene graphs and overall visual semantics to ensure logical coherence and alignment with the images. Extensive user testing demonstrates that our algorithm significantly outperforms existing methods in terms of layout richness and quality. The code and dataset will be available at https://github.com/HiHiAllen/Imaginarium.

3D生成视觉引导场景布局

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。