arXiv:2506.07570cs.CVcs.AI2025-06NeurIPS被引 21

用大模型生成更符合人类审美的室内布局,效果超越现有方法。

OptiScene: LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference Optimization

  • 通过人机协作合成1.7万组真实感强的3D场景数据
  • 分两阶段优化:先生成空间描述,再对齐人类偏好,成功率显著提升
  • 开源模型支持交互编辑与机器人导航,适合设计与AI应用

自动室内布局生成在室内设计、虚拟环境构建和具身智能领域具有重要潜力。现有方法分为两类:依赖专有大模型接口(如GPT API)的提示驱动型方法,以及基于扩散模型训练的学习型方法。前者常出现空间不一致且计算成本高,后者受限于粗粒度关系图与小规模数据集,泛化能力差。本文重新审视基于大模型的布局生成,提出3D-SynthPlace大规模数据集,采用‘GPT生成+人工审核’流程,由3D-Front数据集升级而来,包含近1.7万个场景,覆盖卧室、客厅、厨房、浴室四种常见房间类型,含多样化物体及高层级空间标注。进一步提出OptiScene,一个面向室内布局生成的开源大模型,基于3D-SynthPlace进行两阶段训练。第一阶段为监督微调(SFT),教会模型先生成高层空间描述,再条件预测具体物体布局;第二阶段采用多轮直接偏好优化(DPO),有效提升布局质量与生成成功率。大量实验表明,OptiScene优于传统提示驱动与学习型基线模型,并在场景编辑与机器人导航等交互任务中展现良好潜力。

原文摘要 · Abstract (English)

Automatic indoor layout generation has attracted increasing attention due to its potential in interior design, virtual environment construction, and embodied AI. Existing methods fall into two categories: prompt-driven approaches that leverage proprietary LLM services (e.g., GPT APIs) and learning-based methods trained on layout data upon diffusion-based models. Prompt-driven methods often suffer from spatial inconsistency and high computational costs, while learning-based methods are typically constrained by coarse relational graphs and limited datasets, restricting their generalization to diverse room categories. In this paper, we revisit LLM-based indoor layout generation and present 3D-SynthPlace, a large-scale dataset that combines synthetic layouts generated via a 'GPT synthesize, Human inspect' pipeline, upgraded from the 3D-Front dataset. 3D-SynthPlace contains nearly 17,000 scenes, covering four common room types -- bedroom, living room, kitchen, and bathroom -- enriched with diverse objects and high-level spatial annotations. We further introduce OptiScene, a strong open-source LLM optimized for indoor layout generation, fine-tuned based on our 3D-SynthPlace dataset through our two-stage training. For the warum-up stage I, we adopt supervised fine-tuning (SFT), which is taught to first generate high-level spatial descriptions then conditionally predict concrete object placements. For the reinforcing stage II, to better align the generated layouts with human design preferences, we apply multi-turn direct preference optimization (DPO), which significantly improving layout quality and generation success rates. Extensive experiments demonstrate that OptiScene outperforms traditional prompt-driven and learning-based baselines. Moreover, OptiScene shows promising potential in interactive tasks such as scene editing and robot navigation.

室内布局大模型生成模型人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。