arXiv:2601.04819cs.AI2026-01被引 6

评测视觉语言模型对建筑图纸的理解能力,发现其识图能力仍不成熟。

AECV-Bench: Benchmarking Multimodal Models on Architectural and Engineering Drawings Understanding

  • 构建建筑图纸理解基准,含120张平面图计数与192个问答任务
  • 文字识别准确率达0.95,但门窗等符号计数仅0.40-0.55
  • 揭示当前模型依赖文本、缺乏空间与符号理解,适合建筑自动化研究者

建筑与工程图纸(AEC)通过符号、布局规范和密集标注编码几何与语义信息,但现有多模态与视觉语言模型是否能可靠解析这种图形语言尚不明确。本文提出AECV-Bench,一个基于真实AEC文件的评估基准,包含两个互补任务:(i) 在120张高质量平面图上进行物体计数(门、窗、卧室、厕所),采用逐字段精确匹配准确率与平均绝对百分比误差(MAPE)衡量;(ii) 跨192个问答对的绘图基础文档问答任务,涵盖文本提取(OCR)、实例计数、空间推理与比较推理,使用大模型作为评判者与人工复核边缘案例。在统一协议下评估一系列先进模型,结果呈现稳定的能力梯度:文本识别与文本导向问答表现最佳(最高0.95准确率),空间推理中等,而符号导向的图纸理解——特别是门和窗的可靠计数——仍属难题(通常0.40-0.55准确率),存在显著比例误差。表明当前系统仅能作为文档助手,缺乏稳健的绘图理解能力,亟需领域专用表示与人机协同工作流以实现高效建筑自动化。

原文摘要 · Abstract (English)

AEC drawings encode geometry and semantics through symbols, layout conventions, and dense annotation, yet it remains unclear whether modern multimodal and vision-language models can reliably interpret this graphical language. We present AECV-Bench, a benchmark for evaluating multimodal and vision-language models on realistic AEC artefacts via two complementary use cases: (i) object counting on 120 high-quality floor plans (doors, windows, bedrooms, toilets), and (ii) drawing-grounded document QA spanning 192 question-answer pairs that test text extraction (OCR), instance counting, spatial reasoning, and comparative reasoning over common drawing regions. Object-counting performance is reported using per-field exact-match accuracy and MAPE results, while document-QA performance is reported using overall accuracy and per-category breakdowns with an LLM-as-a-judge scoring pipeline and targeted human adjudication for edge cases. Evaluating a broad set of state-of-the-art models under a unified protocol, we observe a stable capability gradient; OCR and text-centric document QA are strongest (up to 0.95 accuracy), spatial reasoning is moderate, and symbol-centric drawing understanding - especially reliable counting of doors and windows - remains unsolved (often 0.40-0.55 accuracy) with substantial proportional errors. These results suggest that current systems function well as document assistants but lack robust drawing literacy, motivating domain-specific representations and tool-augmented, human-in-the-loop workflows for an efficient AEC automation.

建筑图纸多模态评测基准视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。