arXiv:2410.17051cs.CL2024-10EMNLP被引 2

从3000万篇生物医学文献中挖掘共指关系,构建数据驱动的本体

Data-driven Coreference-based Ontology Building

  • 基于共指链构建词串图谱,通过中心性分析识别层级、等价与噪声
  • 从3000万篇摘要中提取共指链,生成与人工本体高度重合的领域本体
  • 适合生物信息学、知识图谱研究者使用,开源代码与数据

传统共指消解用于单文档理解,本文从全局视角出发,探索从大规模语料中所有文档级共指关系可获得的领域知识。我们从3000万篇生物医学摘要中提取共指链,构建以链内词串为节点的图结构,若词串同现于同一共指链则建立连接。利用图结构与介数中心性,区分层级、等价与噪声边,为层级边赋予方向,并拆分表示多个概念的节点。最终得到一个丰富且数据驱动的生物医学领域本体,其部分与人工构建本体高度重合。我们以知识共享许可协议发布共指链及所得本体,附带源代码。

原文摘要 · Abstract (English)

While coreference resolution is traditionally used as a component in individual document understanding, in this work we take a more global view and explore what can we learn about a domain from the set of all document-level coreference relations that are present in a large corpus. We derive coreference chains from a corpus of 30 million biomedical abstracts and construct a graph based on the string phrases within these chains, establishing connections between phrases if they co-occur within the same coreference chain. We then use the graph structure and the betweeness centrality measure to distinguish between edges denoting hierarchy, identity and noise, assign directionality to edges denoting hierarchy, and split nodes (strings) that correspond to multiple distinct concepts. The result is a rich, data-driven ontology over concepts in the biomedical domain, parts of which overlaps significantly with human-authored ontologies. We release the coreference chains and resulting ontology under a creative-commons license, along with the code.

本体构建共指消解生物医学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。