构建施工估价专用评测数据集,检验大模型在工程图纸理解中的表现
CEQuest: Benchmarking Large Language Models for Construction Estimation
- 设计专用于施工估价的问答基准数据集
- 五款主流大模型在准确率上均有明显提升空间
- 适合建筑智能化、AI辅助设计方向研究者参考
大语言模型在通用任务中表现出色,但在建筑等专业领域应用仍不充分。本文提出CEQuest,一个专为评估大模型在施工图纸解读与工程估价任务中表现而设计的新基准数据集。我们使用Gemma 3、Phi4、LLaVA、Llama 3.3和GPT-4.1五款先进大模型进行系统实验,从准确率、推理时间与模型规模三方面评估其性能。结果表明,当前大模型在该领域仍有显著改进空间,凸显融入领域知识的重要性。为推动后续研究,我们将开源CEQuest数据集,助力开发面向建筑领域的专用大模型。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of general-domain tasks. However, their effectiveness in specialized fields, such as construction, remains underexplored. In this paper, we introduce CEQuest, a novel benchmark dataset specifically designed to evaluate the performance of LLMs in answering construction-related questions, particularly in the areas of construction drawing interpretation and estimation. We conduct comprehensive experiments using five state-of-the-art LLMs, including Gemma 3, Phi4, LLaVA, Llama 3.3, and GPT-4.1, and evaluate their performance in terms of accuracy, execution time, and model size. Our experimental results demonstrate that current LLMs exhibit considerable room for improvement, highlighting the importance of integrating domain-specific knowledge into these models. To facilitate further research, we will open-source the proposed CEQuest dataset, aiming to foster the development of specialized large language models (LLMs) tailored to the construction domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。