首个面向建筑图纸的多层级图文推理基准,助力AI理解真实工程图。
DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings
- 构建真实施工图多层级问答数据集,覆盖感知、上下文与专家级推理
- 评估顶尖多模态大模型在工程任务中表现,发现高阶推理能力显著不足
- 首次将工程工作流映射至AI推理能力,适合研究工程AI与多模态模型者
我们提出DrawingVQA,首个专为评估多模态大语言模型(MLLMs)在真实建筑图纸上的多层级图文推理能力而设计的基准。与自然图像或简略平面图不同,建筑图纸融合抽象几何、符号标注、表格数据、注释及领域文本,构成工程实践中的核心复杂图文域。DrawingVQA包含33张“已签发用于施工”的图纸和92对专家精心构建的问答对,涵盖三类推理深度:感知理解、上下文解释与领域专家推理。为系统评估模型能力,我们提出双维度分类框架,联合分析七项建筑工程与四项MLLM能力维度,首次明确将工程工作流映射至AI推理能力。对主流MLLM的评估显示,模型表现与专家水平存在显著差距,尤其在高阶推理层面。该基准为专业化多模态推理奠定基础,推动AI理解与真实工程流程的深度融合。
原文摘要 · Abstract (English)
We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings -- a core media in architecture, civil, and many other engineering practices. Unlike natural images or schematic floor plans, construction drawings fuse abstract geometry, symbolic notation, tabular data, annotations, and domain-specific text, forming a uniquely complex visual-textual domain core to engineering workflows. DrawingVQA bridges this gap with 33 "Issued for Construction" drawings and 92 expertly curated question-answer pairs, spanning three reasoning depths: perceptual understanding, contextual interpretation, and domain-expert reasoning. To evaluate model capabilities, we present a dual categorization framework to jointly analyze performance across seven construction-engineering and four MLLM capability dimensions -- the first to explicitly map engineering workflows to AI reasoning competencies. Evaluations of state-of-the-art MLLMs reveal a substantial gap between model and expert performance, particularly at higher reasoning depths. This benchmark lays a foundation for domain-specialized multimodal reasoning to allow for advancement on integration of AI-driven understanding and real-world engineering workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。