评测大模型对建筑信息模型的自然语言编辑能力,发现当前水平远未达标。
BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling

- 构建BIM-Edit基准,测试大模型对IFC格式建筑模型的语义化编辑能力。
- 最佳模型平均得分仅49.5%,超96%任务未能完全解决。
- 聚焦几何、语义与拓扑三维度评估,填补现有设计类评测空白。
大型语言模型(LLMs)正被用于计算机辅助设计(CAD),从文本指令生成设计成果。但在工程实践中,不仅需创建新几何,更需理解现有场景、正确编辑并保持语义与关系一致性。然而,多数现有CAD基准侧重生成新模型,且仅评估几何正确性。本文提出BIM-Edit,一个面向以工业基础类(IFC)格式表示的建筑信息模型(BIM)的自然语言编辑评估基准。BIM因其同时编码几何、语义与关系结构而构成挑战性测试环境。BIM-Edit包含324项编辑任务,覆盖11个真实建筑模型和36个合成场景,指令涵盖直接、空间与拓扑三类,兼顾显式与场景依赖型编辑。评估维度包括几何精度、语义有效性与拓扑一致性。在所测LLMs中,表现最佳者三项指标平均得分仅49.5%,无模型能完全解决超过3.4%的任务。结果表明,当前大模型能力与结构化工程设计需求间存在显著差距。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly applied to computer-aided design (CAD) to generate design artifacts from textual instructions. In engineering practice, this requires more than creating new geometry, models must also understand existing scenes, edit them correctly, and preserve semantics and relations. However, many CAD benchmarks focus on creating new models rather than editing existing ones, and mostly evaluate geometric correctness. We introduce BIM-Edit, a benchmark for evaluating LLMs on natural-language editing of Building Information Models (BIM) represented in the Industry Foundation Classes (IFC) format. BIM provides a challenging testbed because building models encode geometry together with semantic and relational structure. BIM-Edit contains 324 editing tasks spanning 11 realistic building models and 36 synthetic scenes. Tasks are expressed using three instruction categories - direct, spatial, and topological - covering both explicit and scene-grounded edits. We evaluate outputs along three dimensions: geometric accuracy, semantic validity, and topological consistency. Across evaluated LLMs, the best-performing model achieves only 49.5% average score across the three metrics, and no model fully solves more than 3.4% of tasks. These results demonstrate a substantial gap between current LLM capabilities and the requirements of structured engineering design workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。