arXiv:2512.11234cs.CV2025-12

通过多模态输入生成可精准控制的室内场景

RoomPilot: Controllable Indoor Scene Synthesis via Multimodal Semantic Parsing

  • 将文本和图纸转为结构化语义语言IDSL
  • 分层生成建筑-房间-物体,保持布局一致性
  • 适合游戏、设计与具身智能领域应用

可控室内场景生成在游戏开发、建筑可视化与具身AI中至关重要。现有方法或输入模态有限,或依赖隐式生成过程,难以精确控制场景结构与语义。为此,我们提出RoomPilot,一个统一框架,支持文本描述与CAD平面图等多模态输入。该框架将异构输入映射为室内领域特定语言(IDSL),作为结构化且可解释的场景语义表征。基于IDSL,RoomPilot构建分层合成流程,逐级组织建筑、房间与物体层级,提升多房间布局的结构连贯性与功能一致性。此外,构建了带丰富语义标注的资产数据集,提升生成场景的视觉真实感与外观一致性。大量实验表明,该方法在多模态理解、细粒度控制、物理合理性与视觉保真度方面均有显著提升,推动可控3D室内场景生成发展。代码与模型将公开。

原文摘要 · Abstract (English)

Generating controllable indoor scenes is fundamental to applications in game development, architectural visualization, and embodied AI. However, existing approaches either support a limited input modalities or rely on implicit generation processes that hinder precise control over scene structure and semantics. To address these limitations, we introduce RoomPilot, a unified framework for controllable indoor scene synthesis from multi-modal inputs, including textual descriptions and CAD floor plans. RoomPilot maps heterogeneous inputs into an Indoor Domain-Specific Language (IDSL), which serves as a structured and interpretable semantic representation for describing indoor scenes. Built upon IDSL, RoomPilot presents a hierarchical synthesis pipeline that progressively organizes scenes at the building, room, and object levels, promoting structural coherence and functional consistency across multi-room layouts. Moreover, RoomPilot constructs a curated asset dataset with rich semantic annotations to support high-quality scene synthesis, improving visual realism and appearance consistency. Extensive experiments demonstrate effective multi-modal understanding, fine-grained controllability in scene generation, and improved physical consistency and visual fidelity, marking a significant step toward controllable 3D indoor scene synthesis. Code and model will be available.

场景生成多模态可控生成3D建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。