arXiv:2606.06390cs.CVcs.AI2026-06

用大模型生成可控制的完整家居布局,支持真实感与仿真应用。

HomeWorld: A Unified Floorplan-to-Furnished Framework for Generating Controllable, Densely Interactive Whole-Home Scenes

论文配图:HomeWorld: A Unified Floorplan-to-Furnished Framework for Generating Controllable, Densely Interactive Whole-Home Scenes
图 1 · 摘自论文原文
  • 分层统一框架:先生成可控户型,再逐步添置家具与小物件。
  • 30万真实户型训练大模型,支持细粒度设计调整。
  • 适合机器人仿真、室内设计,开源5000个完整家具场景。

室内场景生成对机器人仿真和现代室内设计至关重要,但复杂布局与稀缺3D数据使基于学习的生成充满挑战。现有方法多依赖手工规则或聚焦孤立任务(如户型生成或单间布置),导致整体场景缺乏全局一致性、真实感与仿真可用性。为此,我们提出一个统一的分层框架,将室内场景合成分解为可控制的阶段:首先构建包含30万条真实住宅户型的大规模数据集,训练大型语言模型以生成全屋户型;结合详细描述与基于K-D树的表示,实现细粒度可控的户型生成。在此基础上,利用图像生成模型从多视角漫游出发绘制家具布局,并在不同承托表面(如柜子、桌子、餐桌)上生成可操作小物件布局,用于具身智能仿真。家具与物品布局过程中,基于视觉语言模型的精修模块迭代纠正位置,3D生成模型支持单个资产灵活替换。最后附加基础物理属性及简单材质与光照设置,完成具身智能使用流程。实验与用户研究显示,该流程生成的室内空间具有更高布局多样性与更强三维设计吸引力,在量化与定性指标上均优于先前方法。项目将同步发布户型数据集与5000个完整家具场景供社区使用。

原文摘要 · Abstract (English)

Indoor scene generation is crucial for robot simulation and modern interior design. However, complex layouts together with scarce 3D scene data make learning-based generation challenging. Existing methods often rely on hand-crafted rules or focus on isolated sub-tasks (e.g., floorplan synthesis or single-room furnishing), producing whole-home scenes that lack global coherence, realism, and simulation readiness. To mitigate these limitations, we propose a unified hierarchical framework that decomposes indoor scene synthesis into controllable stages. First, we curate a large-scale dataset of 300K real residential floorplans to train a large language model for whole-home floorplan generation. With detailed descriptions and a K-D tree-based representation, our method enables fine-grained, controllable whole-home floorplan generation. Building upon the generated whole-home floorplan, we leverage image generation models to draft furniture layouts from multi-level roaming viewpoints, and then generate the layouts of small manipulable objects on different supporting surfaces (e.g., cabinets, desks, and dining tables) for embodied AI simulation. During furniture and object layout generation, a VLM-based refiner iteratively corrects furniture and object placement, and a 3D generative model enables flexible replacement of individual assets. We further attach basic physical attributes and simple surface texture and lighting setups to complete the pipeline for embodied AI use. Experiments and user studies demonstrate that our pipeline produces indoor spaces with greater layout diversity and stronger 3D design appeal, outperforming prior methods on both quantitative and qualitative metrics. Finally, alongside our generation pipeline, we will release the floorplan dataset and 5K fully furnished scenes to the community. Project Page: https://kairos-homeworld.github.io/

场景生成具身智能3D布局大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。