arXiv:2509.18776cs.CLcs.AI2025-09中稿 · Advanced Engineeri…被引 10

为建筑领域大模型设计分级评估基准,揭示其在专业场景下的真实能力短板。

AECBench: A Hierarchical Benchmark for Knowledge Evaluation of Large Language Models in the AEC Field

  • 构建五层认知评估框架,覆盖从记忆到应用的完整知识链条。
  • 测试9个模型发现:虽能答基础题,但在复杂推理与文档生成上表现差。
  • 用工程师实测数据+专家评审,确保评估真实可信,适合工程界参考。

大型语言模型(LLMs)在建筑、工程与施工(AEC)领域正获得越来越多的应用,展现出优化建筑全生命周期流程的潜力。然而,其在这一专业化且关乎安全的领域中的鲁棒性与可靠性仍需评估。为此,本文提出AECBench,一个全面的基准评测体系,用于量化当前LLMs在AEC领域的优劣。该基准采用五层级的认知导向评估框架(即知识记忆、理解、推理、计算与应用),并据此定义了23项代表性任务,涵盖从规范检索到专业文档生成的真实工程实践。基于此,由工程师主导构建了一个包含4,800道题目的数据集,涵盖开放问答等多种格式,并经过两轮专家评审验证。此外,引入“以LLM为裁判”方法,利用专家制定的评分标准,实现对长文本回复的可扩展、一致化评估。对9个LLMs的评估显示,模型在知识记忆与理解层面表现尚可,但在解读建筑规范中的表格信息、执行复杂推理与计算、生成领域专用文档方面存在显著不足。本研究为未来在高安全要求工程实践中可靠集成LLMs奠定了基础。

原文摘要 · Abstract (English)

Large language models (LLMs), as a novel information technology, are seeing increasing adoption in the Architecture, Engineering, and Construction (AEC) field. They have shown their potential to streamline processes throughout the building lifecycle. However, the robustness and reliability of LLMs in such a specialized and safety-critical domain remain to be evaluated. To address this challenge, this paper establishes AECBench, a comprehensive benchmark designed to quantify the strengths and limitations of current LLMs in the AEC domain. The benchmark features a five-level, cognition-oriented evaluation framework (i.e., Knowledge Memorization, Understanding, Reasoning, Calculation, and Application). Based on the framework, 23 representative evaluation tasks were defined. These tasks were derived from authentic AEC practice, with scope ranging from codes retrieval to specialized documents generation. Subsequently, a 4,800-question dataset encompassing diverse formats, including open-ended questions, was crafted primarily by engineers and validated through a two-round expert review. Furthermore, an "LLM-as-a-Judge" approach was introduced to provide a scalable and consistent methodology for evaluating complex, long-form responses leveraging expert-derived rubrics. Through the evaluation of nine LLMs, a clear performance decline across five cognitive levels was revealed. Despite demonstrating proficiency in foundational tasks at the Knowledge Memorization and Understanding levels, the models showed significant performance deficits, particularly in interpreting knowledge from tables in building codes, executing complex reasoning and calculation, and generating domain-specific documents. Consequently, this study lays the groundwork for future research and development aimed at the robust and reliable integration of LLMs into safety-critical engineering practices.

大模型评测建筑AI认知评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。