arXiv:2504.08779cs.CLcs.AI2025-04被引 1

AI在建筑管理考试中表现超人,但看图题仍弱

Can AI Master Construction Management (CM)? Benchmarking State-of-the-Art Large Language Models on CM Certification Exams

  • 用真实认证考题构建数据集,测试大模型在建筑管理中的表现
  • GPT-4o和Claude 3.7平均准确率达82%~83%,超过人类及格线
  • 看图题准确率仅约40%,概念理解错误是主要问题

建筑管理项目日益复杂,面临监管严格和人力短缺等挑战,亟需专业分析工具提升效率。尽管大语言模型在通用推理任务中表现优异,其在建筑管理特定任务(如精确数量分析和法规解读)中的有效性尚未充分探索。为此,本研究构建了CMExamSet,一个包含689道来自四类国家级建筑管理认证考试的多选题数据集。零样本评估涵盖整体准确率、主题领域(如施工安全)、推理复杂度(单步与多步)及题型(纯文本、图文参考、表格参考)。结果表明,GPT-4o和Claude 3.7平均准确率分别为82%和83%,超过人类及格线(70%)。两者在单步任务中表现更好,准确率分别为85.7%和86.7%;多步任务降至76.5%和77.6%。此外,图文参考题准确率均降至约40%。错误模式分析显示,概念误解占比最高(分别为44.4%和47.9%),凸显对领域专用推理模型的需求。研究证实大模型可作为建筑管理的辅助分析工具,但需针对性优化并保持人工监督。

原文摘要 · Abstract (English)

The growing complexity of construction management (CM) projects, coupled with challenges such as strict regulatory requirements and labor shortages, requires specialized analytical tools that streamline project workflow and enhance performance. Although large language models (LLMs) have demonstrated exceptional performance in general reasoning tasks, their effectiveness in tackling CM-specific challenges, such as precise quantitative analysis and regulatory interpretation, remains inadequately explored. To bridge this gap, this study introduces CMExamSet, a comprehensive benchmarking dataset comprising 689 authentic multiple-choice questions sourced from four nationally accredited CM certification exams. Our zero-shot evaluation assesses overall accuracy, subject areas (e.g., construction safety), reasoning complexity (single-step and multi-step), and question formats (text-only, figure-referenced, and table-referenced). The results indicate that GPT-4o and Claude 3.7 surpass typical human pass thresholds (70%), with average accuracies of 82% and 83%, respectively. Additionally, both models performed better on single-step tasks, with accuracies of 85.7% (GPT-4o) and 86.7% (Claude 3.7). Multi-step tasks were more challenging, reducing performance to 76.5% and 77.6%, respectively. Furthermore, both LLMs show significant limitations on figure-referenced questions, with accuracies dropping to approximately 40%. Our error pattern analysis further reveals that conceptual misunderstandings are the most common (44.4% and 47.9%), underscoring the need for enhanced domain-specific reasoning models. These findings underscore the potential of LLMs as valuable supplementary analytical tools in CM, while highlighting the need for domain-specific refinements and sustained human oversight in complex decision making.

建筑管理大模型评测AI辅助决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。