提出RKEN框架,提升大规模医学知识图谱的构建精度与泛化能力。
Representation-Enhanced Neural Knowledge Integration with Application to Large-Scale Medical Ontology Learning
- 融合表示学习与神经网络,统一建模多种关系类型。
- 理论证明模型在样本内和样本外误差有界,支持大规模应用。
- 适用于医学知识图谱构建,尤其适合异构关系数据融合。
大规模知识图谱通过提供标准化、集成化的框架,提升了生物医学数据发现的可复现性,并增强研究结果在不同人群和条件下的泛化能力。然而,利用现有文献中的多源信息生成可靠的知识图谱,尤其是在节点数量庞大且关系异质的情况下,仍具挑战性。本文提出一种通用且具有理论保障的统计框架RENKI,实现多种关系类型的联合学习。RENKI可推广统计学与计算机科学中广泛使用的各类网络模型。该框架将表示学习的输出作为神经网络初始实体嵌入,用于近似知识图谱的得分函数,并持续训练以拟合观测事实。我们证明了知识图谱函数类伪维数下,样本内与样本外加权均方误差的非渐近界。此外,给出了基于ReLU激活函数的多层神经网络得分函数在嵌入参数固定或可训练情况下的伪维数。最后,通过数值实验验证理论结果,并将方法应用于整合预训练语言模型表示与多个医学本体中的知识图谱链接,构建综合性医学知识图谱。实验支持理论结论,展示了加权机制在异构关系中的有效性,以及表示学习对非参数模型的增益。
原文摘要 · Abstract (English)
A large-scale knowledge graph enhances reproducibility in biomedical data discovery by providing a standardized, integrated framework that ensures consistent interpretation across diverse datasets. It improves generalizability by connecting data from various sources, enabling broader applicability of findings across different populations and conditions. Generating reliable knowledge graph, leveraging multi-source information from existing literature, however, is challenging especially with a large number of node sizes and heterogeneous relations. In this paper, we propose a general theoretically guaranteed statistical framework, called RENKI, to enable simultaneous learning of multiple relation types. RENKI generalizes various network models widely used in statistics and computer science. The proposed framework incorporates representation learning output into initial entity embedding of a neural network that approximates the score function for the knowledge graph and continuously trains the model to fit observed facts. We prove nonasymptotic bounds for in-sample and out-of-sample weighted MSEs in relation to the pseudo-dimension of the knowledge graph function class. Additionally, we provide pseudo-dimensions for score functions based on multilayer neural networks with ReLU activation function, in the scenarios when the embedding parameters either fixed or trainable. Finally, we complement our theoretical results with numerical studies and apply the method to learn a comprehensive medical knowledge graph combining a pretrained language model representation with knowledge graph links observed in several medical ontologies. The experiments justify our theoretical findings and demonstrate the effect of weighting in the presence of heterogeneous relations and the benefit of incorporating representation learning in nonparametric models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。