arXiv:2603.11640cs.CVcs.AI2026-03中稿 · CVPR

用离散房间令牌让大模型读懂、生成和编辑建筑平面图

Tokenization Allows Multimodal Large Language Models to Understand, Generate and Edit Architectural Floor Plans

  • 用离散房间实例令牌构建统一词汇表,连接布局与符号推理
  • 根据文本指令生成符合几何合理性的可控平面图,效率高且可本地部署
  • 适合建筑生成、智能设计助手等场景,支持理解、生成与编辑一体化

建筑平面图设计需要对几何、语义和空间层级进行联合推理,这对现有AI系统仍是重大挑战。尽管近期扩散模型和语言模型提升了视觉保真度,但在连贯的空间推理和可控生成方面仍存在不足。我们提出HouseMind,一个统一建筑平面图理解、生成与编辑的多模态大模型。引入离散房间实例令牌构建统一词汇表,实现布局与符号推理的桥梁。通过多模态对齐与指令微调,模型可从文本指令合成连贯、可控的布局。实验表明,该框架在保持高效性的同时,显著提升几何合理性与可控性,且支持本地部署。

原文摘要 · Abstract (English)

Architectural floor plan design demands joint reasoning over geometry, semantics, and spatial hierarchy, which remains a major challenge for current AI systems. Although recent diffusion and language models improve visual fidelity, they still struggle with coherent spatial reasoning and controllable generation. We present HouseMind, a multimodal large language model that unifies floor plan understanding, generation, and editing in one framework. We introduce discrete room-instance tokens to construct a unified vocabulary that bridges layouts and symbolic reasoning. With multimodal alignment and instruction tuning, the model synthesizes coherent, controllable layouts from text instructions. Experiments show how the framework achieves superior geometric validity and controllability while remaining efficient and locally deployable.

建筑生成多模态可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。