用AI从法语建筑规范中自动提取结构化需求,提升BIM建模效率
Extracting Structured Requirements from Unstructured Building Technical Specifications for Building Information Modeling
- 基于CamemBERT和Fr_core_news_lg模型做命名实体识别
- 命名实体识别F1超90%,关系抽取用随机森林达80%以上
- 适合建筑信息化、BIM自动化领域研究者参考
本研究探索将建筑信息模型(BIM)与自然语言处理(NLP)结合,自动化提取建筑行业非结构化法语建筑技术规范(BTS)文档中的要求。采用命名实体识别(NER)和关系抽取(RE)技术,使用基于Transformer的CamemBERT模型,并通过在通用领域大规模法语文本上预训练的Fr_core_news_lg模型进行迁移学习。为评估性能,构建了从规则方法到深度学习方法的多种对比方案。在关系抽取方面,实现四种监督模型,包括随机森林,使用自定义特征向量。利用人工标注的数据集对比不同NER方法与RE模型的效果。结果表明,CamemBERT与Fr_core_news_lg在NER任务中表现优异,F1分数超过90%;随机森林在关系抽取中最为有效,F1分数高于80%。未来工作计划将结果以知识图谱形式表示,进一步增强自动验证系统。
原文摘要 · Abstract (English)
This study explores the integration of Building Information Modeling (BIM) with Natural Language Processing (NLP) to automate the extraction of requirements from unstructured French Building Technical Specification (BTS) documents within the construction industry. Employing Named Entity Recognition (NER) and Relation Extraction (RE) techniques, the study leverages the transformer-based model CamemBERT and applies transfer learning with the French language model Fr\_core\_news\_lg, both pre-trained on a large French corpus in the general domain. To benchmark these models, additional approaches ranging from rule-based to deep learning-based methods are developed. For RE, four different supervised models, including Random Forest, are implemented using a custom feature vector. A hand-crafted annotated dataset is used to compare the effectiveness of NER approaches and RE models. Results indicate that CamemBERT and Fr\_core\_news\_lg exhibited superior performance in NER, achieving F1-scores over 90\%, while Random Forest proved most effective in RE, with an F1 score above 80\%. The outcomes are intended to be represented as a knowledge graph in future work to further enhance automatic verification systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。