arXiv:2605.22079cs.CL2026-05中稿 · CIKM '26被引 1

首个面向BIM信息交付规范生成的公开基准,评估大模型在专业术语和格式约束下的表现。

Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements

论文配图:Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
图 1 · 摘自论文原文
  • 构建首个公开可用的IDS生成基准,涵盖83个实际场景与双语专家标注数据。
  • 10个大模型零样本测试中最高内容准确率仅65.6%,格式通过率不足33.1%。
  • 适合研究建筑信息模型、AI辅助设计及标准合规性验证的开发者与学者。

建筑信息模型(BIM)项目越来越多地使用信息交付规范(IDS)以机器可读的XML格式形式化信息需求。由于IDS条件基于工业基础类(IFC)词汇,编写需精通IFC概念、验证工具及属性集规范。现有结构化生成基准未能充分反映IDS对词汇一致性与外部验证器兼容性的额外要求。本文提出Ishigaki-IDS-Bench,首个公开发布的从BIM信息需求生成IDS的基准。该基准包含166个案例,覆盖83个实际场景,由六位BIM/IDS专家以日文和英文编写,每例均配有黄金标准IDS文件及元数据,涵盖输入格式、交互设置、目标IFC版本和建筑领域。评估分两阶段进行:(i) 通过buildingSMART IDSAuditTool对可处理性、结构和内容进行形式有效性评分;(ii) 以面级宏F1与黄金标准对比内容保真度。在零样本条件下对10个大模型测试,最高面级F1为65.6%(GPT-5.5),最高内容通过率仅为33.1%(Claude Opus 4.5)。Ishigaki-IDS-Bench已在Hugging Face(DOI 10.57967/hf/8873)以CC BY 4.0发布,评估代码在Zenodo(DOI 10.5281/zenodo.20550510)以Apache-2.0发布。

原文摘要 · Abstract (English)

Building Information Modeling (BIM) projects increasingly use Information Delivery Specification (IDS) to formalize information requirements in a machine-checkable XML format. Because IDS conditions are grounded in the Industry Foundation Classes (IFC) vocabulary, authoring them requires expertise in IFC concepts, validation tools, and property set conventions. Existing benchmarks for structured generation do not adequately capture the additional burden of vocabulary conformance and external-validator agreement that IDS imposes. We present Ishigaki-IDS-Bench, the first publicly released benchmark for IDS generation from BIM information requirements. The benchmark contains 166 examples spanning 83 practical scenarios authored in Japanese and English by six BIM/IDS experts, each paired with a gold IDS file and metadata covering input format, turn setting, target IFC versions, and construction domain. Evaluation proceeds in two stages: (i) formal validity scored by the buildingSMART IDSAuditTool along Processability, Structure, and Content, and (ii) content fidelity scored by facet-level macro-F1 against the gold IDS. Across 10 LLMs in zero-shot, the highest Facet F1 is 65.6%, achieved by GPT-5.5, while the highest Content pass rate is only 33.1%, achieved by Claude Opus 4.5. Ishigaki-IDS-Bench is released on Hugging Face (DOI 10.57967/hf/8873) under CC BY 4.0, and the evaluation code is released on Zenodo (DOI 10.5281/zenodo.20550510) under Apache-2.0.

BIMIDS生成大模型评测建筑信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。