arXiv:2608.09281cs.AI2026-08

测试大模型能否结合建筑图与工程原理推理,发现现有模型差距巨大。

MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence

论文配图:MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence
图 1 · 摘自论文原文
  • 用论文图表构建跨领域推理题,避免单图或文本捷径。
  • 最强开源模型仅30%正确,顶尖闭源系统52%,人类专家达95%。
  • 失败主因是整合多图证据和应用原理,适合建筑AI研究者参考。

多模态大语言模型在工程图像任务中表现优异,但现有基准主要测试绘图识别、信息提取或合规检查,未检验模型是否能结合分散的视觉证据与工程原理得出结论。我们提出MMArch,一个涵盖十个子领域的建筑与土木工程基准,全部基于同行评审论文中的图表构建。其1,212道简答题通过解耦的规划-撰写流程生成,并经自动化筛选、盲式对抗审计和专家评审验证,确保回答需感知相关证据、识别核心原理并加以应用,而非依赖文本或单图线索。对18个开源与专有模型的评估显示,最强开源模型准确率约30%,最佳闭源系统为52%,而人类专家达到95%,差距超过40个百分点。错误分析表明,失败主要集中在原理应用和跨图证据整合,而非证据定位,提示未来研究仍有巨大提升空间。代码与数据已公开于https://dcx-swjtu.github.io/MMArch/。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information extraction, or compliance checking, leaving open whether models can combine distributed visual evidence with engineering principles to reach a conclusion. We introduce MMArch, a benchmark for architecture and civil engineering spanning ten subdomains and built entirely from figures in peer-reviewed papers. Its $1{,}212$ short-answer items are produced by a decoupled planner--writer pipeline and validated through automated screening, a blind adversarial audit, and expert review, so that answering requires perceiving the relevant evidence, identifying the governing principle, and applying it, not exploiting textual or single-figure shortcuts. Evaluating $18$ open-weight and proprietary MLLMs against a domain-expert panel, we find a wide gap: the strongest open-source model attains about $30\%$ and the best proprietary system $52\%$, while human experts reach $95\%$, more than forty points ahead. Our error analysis shows that failures concentrate in applying principles and combining evidence across figures rather than in locating it, pointing to substantial headroom for future research. Code and data are available at https://dcx-swjtu.github.io/MMArch/.

多模态推理建筑AI评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。