用深度学习从建筑文本中自动提炼术语与上下位关系。
Deep Learning and Natural Language Processing in the Field of Construction
- 先用统计和n-gram提取建筑术语,再通过语言模式和网络查询优化。
- 结合多种词向量模型,准确识别术语间的上下位关系。
- 经专家评估验证,适合建筑知识图谱构建与智能检索。
本文提出一种完整的建筑领域超义关系抽取流程,包含两个主要步骤:术语提取与上下位关系检测。首先,通过语料库分析方法,利用统计和词n-gram分析从建筑技术规范文本中提取领域术语,并结合语言模式和网络查询进行筛选优化,提升术语质量。其次,采用基于多种词嵌入模型及组合的机器学习方法,实现从提取术语中识别上下位关系。术语提取结果由6名领域专家进行人工评估,上下位关系检测在多个数据集上进行了测试。整体方法表现良好,结果具有显著实用价值。
原文摘要 · Abstract (English)
This article presents a complete process to extract hypernym relationships in the field of construction using two main steps: terminology extraction and detection of hypernyms from these terms. We first describe the corpus analysis method to extract terminology from a collection of technical specifications in the field of construction. Using statistics and word n-grams analysis, we extract the domain's terminology and then perform pruning steps with linguistic patterns and internet queries to improve the quality of the final terminology. Second, we present a machine-learning approach based on various words embedding models and combinations to deal with the detection of hypernyms from the extracted terminology. Extracted terminology is evaluated using a manual evaluation carried out by 6 experts in the domain, and the hypernym identification method is evaluated with different datasets. The global approach provides relevant and promising results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。