arXiv:2608.24886cs.AI2026-08

用视觉语言模型自动构建建筑平面图的多粒度图表示,提升设计信息利用效率。

VLM-based automatic multi-granularity graph representation of building layouts for design informatics

论文配图:VLM-based automatic multi-granularity graph representation of building layouts for design informatics
图 1 · 摘自论文原文
  • 基于VLM模型,通过节点识别、边推断等步骤自动构建多粒度图结构。
  • 生成的图与人工标注高度一致(节点匹配率≥92%),三粒度图每张图生成仅需509.3秒。
  • 中粒度图在区域预测上表现最优,粗粒度图适合布局质量评估,适合设计检索与BIM应用。

建筑平面图蕴含功能空间间的丰富关系知识,支撑设计检索、基于知识的推理及建筑全生命周期的BIM增强。然而,自动构建适用于公共建筑的任务自适应图表示仍具挑战。为此,我们首次定义了公共建筑布局的多粒度图层次(LoGs)。方法上,提出基于视觉语言模型(VLM)的自动LoG构建框架,包括节点识别、边推理、文本解析和图粗化。在147个全球学术图书馆平面图的案例研究中,系统评估了VLM生成的图表示在真实任务中的表现。实验表明,其与人工标注结果高度一致(节点匹配率≥92%;每张平面图生成三粒度图耗时509.3秒)。中粒度图在区域预测任务中表现最佳(宏观F1=0.647,复杂度仅为细粒度的65%),而粗粒度图在布局质量评估上效果最优(斯皮尔曼相关系数ρ=0.610,复杂度仅为细粒度的16%)。该方法实现了从平面图中可扩展、免标注地提取结构化布局信息,推动设计信息化发展,促进建筑全生命周期中设计信息的有效利用。

原文摘要 · Abstract (English)

Architectural floorplan images encode rich relational knowledge among functional spaces, which underpins design retrieval, knowledge-based reasoning, and BIM enrichment through the building lifecycle. However, it remains challenging to automatically construct task-adaptive graph representations for public buildings. To address this gap, we first define a multi-granularity Level-of-Graphs (LoGs) for public building layouts. Methodologically, we present a Vision-Language Model (VLM)-based automatic LoG construction through node identification, edge inference, text parsing, and graph coarsening. VLM-generated representations are systematically evaluated and tested in real-world tasks, using 147 academic library floorplans worldwide as a case study. Experiments showed VLM-generated graphs were broadly consistent with human-labeled graphs (matched node ratio >= 92%; 509.3 s per floor plan for three-LoG graph generation). Meso-grained graphs yield the best node-level zone prediction (Macro F1 = 0.647, at 65% of fine-grained complexity), while coarse-grained graphs are most effective for graph-level layout quality evaluation (Spearman's \r{ho} = 0.610, at 16% of fine-grained complexity). By enabling scalable, annotation-free extraction of structured layout information from floorplan images, this study advances design informatics by converting plan images into knowledge representations, thereby enhancing the utilization of design information across the building life cycle.

建筑信息模型视觉语言模型设计信息化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。