评测建筑领域智能体在真实任务中的表现,推动行业AI应用落地。
AEC-Bench: A Multimodal Benchmark for Agentic Systems in Architecture, Engineering, and Construction
- 构建涵盖图纸理解、跨图推理与项目协调的多模态评估体系。
- 发现通用工具和设计技巧可提升多个基础模型的性能表现。
- 开源数据集与代码,支持研究复现与行业应用探索。
AEC-Bench 是一个针对建筑、工程与施工(AEC)领域真实任务的多模态基准测试平台,涵盖需图纸理解、跨图纸推理及项目级协作的任务。本报告介绍其设计动机、数据集分类体系、评估协议,并提供多个领域专用基础模型的基线结果。通过 AEC-Bench,我们识别出在不同基础模型(如 Claude Code、Codex)的原生调用框架中均能持续提升性能的通用工具与架构设计方法。相关数据集、智能体调用框架与评估代码已开放发布于 https://github.com/nomic-ai/aec-bench,采用 Apache 2 许可证,确保研究可复现性。
原文摘要 · Abstract (English)
The AEC-Bench is a multimodal benchmark for evaluating agentic systems on real-world tasks in the Architecture, Engineering, and Construction (AEC) domain. The benchmark covers tasks requiring drawing understanding, cross-sheet reasoning, and construction project-level coordination. This report describes the benchmark motivation, dataset taxonomy, evaluation protocol, and baseline results across several domain-specific foundation model harnesses. We use AEC-Bench to identify consistent tools and harness design techniques that uniformly improve performance across foundation models in their own base harnesses, such as Claude Code and Codex. We openly release our benchmark dataset, agent harness, and evaluation code for full replicability at https://github.com/nomic-ai/aec-bench under an Apache 2 license.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。