arXiv:2602.15791cs.AIcs.CL2026-02被引 1

用大模型嵌入提升建筑语义区分度,让AI更懂建筑细节

Enhancing Building Semantics Preservation in AI Model Training with Large Language Model Encodings

  • 用大语言模型生成嵌入向量替代传统编码,保留建筑子类间细微差异
  • 在5个高层住宅BIM数据上,最高达0.8766的加权F1分数,优于一维编码
  • 适合建筑信息化、智能设计等领域的语义理解任务

准确表征建筑语义(包括通用类别和具体子类)对建筑、工程、施工与运营(AECO)领域的人工智能模型训练至关重要。传统编码方法(如一维编码)难以传达密切相关的子类间的细微关系,限制了AI的语义理解能力。为此,本研究提出一种新训练方法,采用大语言模型(如OpenAI GPT和Meta LLaMA)的嵌入作为编码,以保留建筑语义的精细差异。通过在五个高层住宅建筑信息模型(BIMs)上训练GraphSAGE模型,对42种建筑对象子类进行分类,测试了不同维度的嵌入,包括原始高维嵌入(1,536、3,072或4,096)及通过马特里什卡表示模型压缩得到的1,024维嵌入。实验表明,大模型编码显著优于传统一维编码基线:llama-3(压缩版)嵌入达到0.8766的加权平均F1分数,而一维编码为0.8475。结果证明,利用大模型嵌入可有效提升AI对复杂领域语义的理解能力。随着大模型与降维技术的发展,该方法在AECO行业语义细化任务中具有广泛应用潜力。

原文摘要 · Abstract (English)

Accurate representation of building semantics, encompassing both generic object types and specific subtypes, is essential for effective AI model training in the architecture, engineering, construction, and operation (AECO) industry. Conventional encoding methods (e.g., one-hot) often fail to convey the nuanced relationships among closely related subtypes, limiting AI's semantic comprehension. To address this limitation, this study proposes a novel training approach that employs large language model (LLM) embeddings (e.g., OpenAI GPT and Meta LLaMA) as encodings to preserve finer distinctions in building semantics. We evaluated the proposed method by training GraphSAGE models to classify 42 building object subtypes across five high-rise residential building information models (BIMs). Various embedding dimensions were tested, including original high-dimensional LLM embeddings (1,536, 3,072, or 4,096) and 1,024-dimensional compacted embeddings generated via the Matryoshka representation model. Experimental results demonstrated that LLM encodings outperformed the conventional one-hot baseline, with the llama-3 (compacted) embedding achieving a weighted average F1-score of 0.8766, compared to 0.8475 for one-hot encoding. The results underscore the promise of leveraging LLM-based encodings to enhance AI's ability to interpret complex, domain-specific building semantics. As the capabilities of LLMs and dimensionality reduction techniques continue to evolve, this approach holds considerable potential for broad application in semantic elaboration tasks throughout the AECO industry.

建筑语义大模型编码语义理解人工智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。