arXiv:2412.10054cs.CL2024-12EMNLP被引 1

无需标注数据,用图算法提升低资源领域实体消歧准确率。

Unsupervised Named Entity Disambiguation for Low Resource Domains

  • 基于组斯坦纳树构建候选实体上下文相似图
  • 在多个领域数据集上精度提升超40%(Precision@1)
  • 适合无标注数据、小样本的特定领域应用

在自然语言处理与信息检索不断发展的背景下,构建稳健且领域特定的实体链接算法变得尤为重要。在人文学科、技术写作和生物医学等领域,为文本注入语义并发现更多知识至关重要。这些领域的命名实体消歧(NED)面临噪声文本、低资源设置及领域专用知识库的挑战。现有方法大多不适用,因它们依赖训练数据或无法灵活适配领域知识库。为此,本文提出一种无监督方法,利用组斯坦纳树(Group Steiner Trees, GST)概念,通过文档中所有提及项的候选实体间上下文相似性,识别最相关候选实体。该方法在多个领域特定数据集上,平均精度(Precision@1)超越现有最优无监督方法超过40%。

原文摘要 · Abstract (English)

In the ever-evolving landscape of natural language processing and information retrieval, the need for robust and domain-specific entity linking algorithms has become increasingly apparent. It is crucial in a considerable number of fields such as humanities, technical writing and biomedical sciences to enrich texts with semantics and discover more knowledge. The use of Named Entity Disambiguation (NED) in such domains requires handling noisy texts, low resource settings and domain-specific KBs. Existing approaches are mostly inappropriate for such scenarios, as they either depend on training data or are not flexible enough to work with domain-specific KBs. Thus in this work, we present an unsupervised approach leveraging the concept of Group Steiner Trees (GST), which can identify the most relevant candidates for entity disambiguation using the contextual similarities across candidate entities for all the mentions present in a document. We outperform the state-of-the-art unsupervised methods by more than 40\% (in avg.) in terms of Precision@1 across various domain-specific datasets.

实体消歧无监督学习低资源知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。