用视觉语言模型实现文本驱动的3D场景分层生成与实时编辑
HOG-Layout: Hierarchical 3D Scene Generation, Optimization and Editing via Vision-Language Models

- 通过分层表示和检索增强生成,提升场景语义一致性
- 引入优化模块,使生成场景更符合物理合理性
- 支持快速直观编辑,适合虚拟现实与智能交互应用
3D布局生成与编辑在具身智能和沉浸式虚拟现实交互中至关重要。然而,手动创建耗时费力,数据驱动方法又常缺乏多样性。大模型的出现为3D场景合成带来新可能。本文提出HOG-Layout,利用大语言模型(LLMs)和视觉语言模型(VLMs)实现文本驱动的分层场景生成、优化与实时编辑。通过检索增强生成(RAG)技术提升场景语义一致性,引入优化模块增强物理合理性,并采用分层表示提升推理与优化效率,实现实时编辑。实验表明,相比现有基线方法,HOG-Layout生成的环境更具合理性,同时支持快速直观的场景编辑。
原文摘要 · Abstract (English)
3D layout generation and editing play a crucial role in Embodied AI and immersive VR interaction. However, manual creation requires tedious labor, while data-driven generation often lacks diversity. The emergence of large models introduces new possibilities for 3D scene synthesis. We present HOG-Layout that enables text-driven hierarchical scene generation, optimization and real-time scene editing with large language models (LLMs) and vision-language models (VLMs). HOG-Layout improves scene semantic consistency and plausibility through retrieval-augmented generation (RAG) technology, incorporates an optimization module to enhance physical consistency, and adopts a hierarchical representation to enhance inference and optimization, achieving real-time editing. Experimental results demonstrate that HOG-Layout produces more reasonable environments compared with existing baselines, while supporting fast and intuitive scene editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。