分步生成3D场景,让用户逐步添加关键物体,自动补全周边细节。
HiGS: Hierarchical Generative Scene Framework for Multi-Step Associative Semantic Spatial Composition
- 分层生成框架,从全局到局部逐步构建场景
- 通过语义关联图提升布局合理性和风格一致性
- 适合需要精细控制的虚拟现实与游戏场景设计
三维场景生成在游戏、影视和虚拟现实领域具有重要潜力。然而,现有方法多采用单步生成,难以在复杂度与用户输入量之间取得平衡。受人类认知建模过程启发——从整体到局部、聚焦关键元素、通过语义关联完成场景——我们提出HiGS,一种用于多步关联语义空间组合的分层生成框架。该框架允许用户通过选择关键语义对象,迭代扩展场景,对兴趣区域实现细粒度控制,同时模型自动完成外围区域。为支持结构化且连贯的生成,我们引入渐进式分层时空-语义图(PHiSSG),动态组织演化场景中的空间关系与语义依赖。PHiSSG通过保持图节点与生成对象的一一映射,并支持递归布局优化,确保生成过程中的空间与几何一致性。实验表明,相比单阶段方法,HiGS在布局合理性、风格一致性和用户偏好上表现更优,提供了一种可控且可扩展的高效3D场景构建范式。
原文摘要 · Abstract (English)
Three-dimensional scene generation holds significant potential in gaming, film, and virtual reality. However, most existing methods adopt a single-step generation process, making it difficult to balance scene complexity with minimal user input. Inspired by the human cognitive process in scene modeling, which progresses from global to local, focuses on key elements, and completes the scene through semantic association, we propose HiGS, a hierarchical generative framework for multi-step associative semantic spatial composition. HiGS enables users to iteratively expand scenes by selecting key semantic objects, offering fine-grained control over regions of interest while the model completes peripheral areas automatically. To support structured and coherent generation, we introduce the Progressive Hierarchical Spatial-Semantic Graph (PHiSSG), which dynamically organizes spatial relationships and semantic dependencies across the evolving scene structure. PHiSSG ensures spatial and geometric consistency throughout the generation process by maintaining a one-to-one mapping between graph nodes and generated objects and supporting recursive layout optimization. Experiments demonstrate that HiGS outperforms single-stage methods in layout plausibility, style consistency, and user preference, offering a controllable and extensible paradigm for efficient 3D scene construction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。