arXiv:2411.12056cs.CLcs.AI2024-11被引 2

评测预训练文本嵌入模型在建筑资产信息对齐中的表现

Benchmarking pre-trained text embedding models in aligning built asset information

  • 构建了基于两大建筑分类词典的六组评测数据集
  • 在聚类、检索、重排序任务中验证模型效果
  • 开源资源助力建筑领域语义对齐研究

准确将建筑资产信息映射到既定数据分类体系和术语库,对项目移交合规性及临时数据整合至关重要。由于建筑资产数据以技术文本为主,该过程仍主要依赖人工和领域专家。近年来,通过预训练大语言模型实现的上下文文本表示学习(文本嵌入)为自动化跨映射提供了可能。然而,尚未有全面评估这些模型在建筑领域特定技术术语复杂语义表征方面的能力。本研究对比评估了当前最优的文本嵌入模型在建筑资产信息与领域术语对齐中的有效性。所用数据集源自两个著名的建筑资产分类词典。在六个数据集上,涵盖聚类、检索、重排序三项任务的基准测试结果表明,亟需未来研究关注领域适应技术。评测资源已作为开源库发布,并将持续维护与扩展,以支持该领域的后续评估。

原文摘要 · Abstract (English)

Accurate mapping of the built asset information to established data classification systems and taxonomies is crucial for effective asset management, whether for compliance at project handover or ad-hoc data integration scenarios. Due to the complex nature of built asset data, which predominantly comprises technical text elements, this process remains largely manual and reliant on domain expert input. Recent breakthroughs in contextual text representation learning (text embedding), particularly through pre-trained large language models, offer promising approaches that can facilitate the automation of cross-mapping of the built asset data. However, no comprehensive evaluation has yet been conducted to assess these models' ability to effectively represent the complex semantics specific to built asset technical terminology. This study presents a comparative benchmark of state-of-the-art text embedding models to evaluate their effectiveness in aligning built asset information with domain-specific technical concepts. Our proposed datasets are derived from two renowned built asset data classification dictionaries. The results of our benchmarking across six proposed datasets, covering three tasks of clustering, retrieval, and reranking, highlight the need for future research on domain adaptation techniques. The benchmarking resources are published as an open-source library, which will be maintained and extended to support future evaluations in this field.

文本嵌入建筑信息领域适配评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。