首个评估大模型建筑病害推理能力的分级基准,揭示其诊断强但定位弱。
Can Large Multimodal Models Inspect Buildings? A Hierarchical Benchmark for Structural Pathology Reasoning

- 构建统一层级语义框架,整合12个分散数据集
- 18个SOTA模型在三重认知维度测试中定位精度不足
- 零样本生成分割可行,通用模型可媲美专用网络
自动化建筑立面检测是城市韧性与智慧城市建设的关键。传统方法依赖专用判别模型(如YOLO、Mask R-CNN),虽擅长像素级定位,却受限于被动感知与泛化能力差,缺乏对结构拓扑的理解。大视觉语言模型(LMMs)有望实现主动推理,但在高风险工程领域尚无严格评估标准。为此,我们提出人机协同半自动标注框架,通过专家验证统一12个碎片化数据集,建立标准化层次化本体。基于此,我们构建首个多维评估基准DefectBench,用于检验LMMs超越基础语义识别的能力。DefectBench在三个递增认知维度(语义感知、空间定位、生成几何分割)上评估18个SOTA LMMs。大量实验表明,当前LMMs虽具备出色的拓扑意识和语义理解能力(有效诊断“是什么”和“如何”),但在度量定位精度(“在哪里”)方面存在显著缺陷。然而,我们验证了零样本生成分割的可行性,证明通用基础模型无需领域训练即可媲美专用监督网络。本工作提供了严格的评估标准与高质量开源数据库,为自主AI代理在土木工程中的发展建立了新基线。
原文摘要 · Abstract (English)
Automated building facade inspection is a critical component of urban resilience and smart city maintenance. Traditionally, this field has relied on specialized discriminative models (e.g., YOLO, Mask R-CNN) that excel at pixel-level localization but are constrained to passive perception and worse generization without the visual understandng to interpret structural topology. Large Multimodal Models (LMMs) promise a paradigm shift toward active reasoning, yet their application in such high-stakes engineering domains lacks rigorous evaluation standards. To bridge this gap, we introduce a human-in-the-loop semi-automated annotation framework, leveraging expert-proposal verification to unify 12 fragmented datasets into a standardized, hierarchical ontology. Building on this foundation, we present \textit{DefectBench}, the first multi-dimensional benchmark designed to interrogate LMMs beyond basic semantic recognition. \textit{DefectBench} evaluates 18 state-of-the-art (SOTA) LMMs across three escalating cognitive dimensions: Semantic Perception, Spatial Localization, and Generative Geometry Segmentation. Extensive experiments reveal that while current LMMs demonstrate exceptional topological awareness and semantic understanding (effectively diagnosing "what" and "how"), they exhibit significant deficiencies in metric localization precision ("where"). Crucially, however, we validate the viability of zero-shot generative segmentation, showing that general-purpose foundation models can rival specialized supervised networks without domain-specific training. This work provides both a rigorous benchmarking standard and a high-quality open-source database, establishing a new baseline for the advancement of autonomous AI agents in civil engineering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。