用大模型生成分层场景描述,实现更合理且符合用户需求的室内布局。
Hierarchically-Structured Open-Vocabulary Indoor Scene Synthesis with Pre-trained Large Language Model
- 通过分层结构描述物体位置关系,提升布局合理性
- 采用分治优化算法,有效求解复杂场景布局
- 适合需要灵活定制室内设计的交互式应用
室内场景合成旨在自动生成合理、真实且多样的3D室内场景,尤其针对任意用户需求。近期预训练大语言模型(LLM)的泛化能力为开放词汇场景合成带来希望,但挑战在于将LLM输出转化为合理且物理可行的场景布局。本文提出使用LLM生成分层结构化的场景描述,并据此计算场景布局。具体而言,我们训练了一个层级感知网络以推断物体间的细粒度相对位置,并设计分治优化算法求解场景布局。分层结构的优势有二:一是提供物体排列的粗略依据,缓解密集关系导致的矛盾布局,增强网络对细粒度位置的泛化能力;二是天然支持分治优化,先安排子场景再整合全局场景,更高效求解可行布局。我们在多种定性与定量评估中进行了广泛对比实验与消融研究,验证了分层结构表示的关键设计有效性。所提方法生成的布局更合理,且与用户需求和LLM描述更一致。此外,我们展示了开放词汇场景合成与交互式场景设计结果,体现了该方法在实际应用中的优势。
原文摘要 · Abstract (English)
Indoor scene synthesis aims to automatically produce plausible, realistic and diverse 3D indoor scenes, especially given arbitrary user requirements. Recently, the promising generalization ability of pre-trained large language models (LLM) assist in open-vocabulary indoor scene synthesis. However, the challenge lies in converting the LLM-generated outputs into reasonable and physically feasible scene layouts. In this paper, we propose to generate hierarchically structured scene descriptions with LLM and then compute the scene layouts. Specifically, we train a hierarchy-aware network to infer the fine-grained relative positions between objects and design a divide-and-conquer optimization to solve for scene layouts. The advantages of using hierarchically structured scene representation are two-fold. First, the hierarchical structure provides a rough grounding for object arrangement, which alleviates contradictory placements with dense relations and enhances the generalization ability of the network to infer fine-grained placements. Second, it naturally supports the divide-and-conquer optimization, by first arranging the sub-scenes and then the entire scene, to more effectively solve for a feasible layout. We conduct extensive comparison experiments and ablation studies with both qualitative and quantitative evaluations to validate the effectiveness of our key designs with the hierarchically structured scene representation. Our approach can generate more reasonable scene layouts while better aligned with the user requirements and LLM descriptions. We also present open-vocabulary scene synthesis and interactive scene design results to show the strength of our approach in the applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。