用上下文词嵌入提升领域术语提取效果
Extracting domain-specific terms using contextual word embeddings
- 结合传统特征与上下文嵌入生成新特征
- 在4个领域上F1分数显著优于现有方法
- 适合语言学、生物医学等领域的术语挖掘
自动术语提取旨在从领域文本中识别有意义的术语。本文提出一种新型机器学习方法,将传统术语提取系统的特征与基于上下文词嵌入的新特征相结合。不依赖预定义的词性模式,而是先分析斯洛文尼亚语的标注语料库RSDO5,制定术语候选选择规则,并生成统计、语言学和上下文特征。采用支持向量机训练分类模型,在RSDO5语料库的四个领域(生物力学、语言学、化学、兽医)上进行评估,并与斯洛文尼亚语当前最优方法对比。结果表明,该方法在F1分数上显著优于之前最先进水平,证明上下文词嵌入对提升术语提取有效。
原文摘要 · Abstract (English)
Automated terminology extraction refers to the task of extracting meaningful terms from domain-specific texts. This paper proposes a novel machine learning approach to terminology extraction, which combines features from traditional term extraction systems with novel contextual features derived from contextual word embeddings. Instead of using a predefined list of part-of-speech patterns, we first analyse a new term-annotated corpus RSDO5 for the Slovenian language and devise a set of rules for term candidate selection and then generate statistical, linguistic and context-based features. We use a support-vector machine algorithm to train a classification model, evaluate it on the four domains (biomechanics, linguistics, chemistry, veterinary) of the RSDO5 corpus and compare the results with state-of-art term extraction approaches for the Slovenian language. Our approach provides significant improvements in terms of F1 score over the previous state-of-the-art, which proves that contextual word embeddings are valuable for improving term extraction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。