arXiv:2505.11185cs.LG2025-05

构建高质量生物知识图谱,助力精准医学与药物重定位研究。

VitaGraph: Building a Knowledge Graph for Biologically Relevant Learning Tasks

  • 整合多个公开数据源,清洗并统一生物关系数据
  • 融合分子指纹与基因本体等特征,提升模型表征能力
  • 支持药物重定位、蛋白互作等任务的基准测试

人类生物学的内在复杂性持续带来科学理解挑战。研究人员跨学科协作以拓展对生命关键生物互作的认知。人工智能方法在计算生物学中成为强大工具,图结构能有效建模蛋白质-蛋白质相互作用(PPI)网络和基因功能网络,这些网络是网络医学核心任务的数据基础,如基因-疾病关联预测、药物重定位和多药副作用研究。机器学习模型的可靠预测依赖高质量的基础数据。本文提出一个综合性的多功能生物知识图谱,通过整合并优化多个公开数据集构建。基于药物重定位知识图谱(DRKG),设计了一套流程:a) 清洗DRKG中的不一致与冗余信息;b) 融合主要公开数据源的信息;c) 为图节点丰富表达性特征向量,包括分子指纹与基因本体。生物化学相关特征提升了机器学习模型生成准确且结构良好的嵌入空间的能力。该资源构成一个连贯可靠的生物知识图谱,成为推进计算生物学与精准医学研究的前沿平台。同时,它为图神经网络与网络医学模型提供了相关任务的基准测试机会。我们通过药物重定位、PPI预测与副作用预测三个任务(均建模为链接预测问题)验证了该数据集的有效性。

原文摘要 · Abstract (English)

The intrinsic complexity of human biology presents ongoing challenges to scientific understanding. Researchers collaborate across disciplines to expand our knowledge of the biological interactions that define human life. AI methodologies have emerged as powerful tools across scientific domains, particularly in computational biology, where graph data structures effectively model biological entities such as protein-protein interaction (PPI) networks and gene functional networks. Those networks are used as datasets for paramount network medicine tasks, such as gene-disease association prediction, drug repurposing, and polypharmacy side effect studies. Reliable predictions from machine learning models require high-quality foundational data. In this work, we present a comprehensive multi-purpose biological knowledge graph constructed by integrating and refining multiple publicly available datasets. Building upon the Drug Repurposing Knowledge Graph (DRKG), we define a pipeline tasked with a) cleaning inconsistencies and redundancies present in DRKG, b) coalescing information from the main available public data sources, and c) enriching the graph nodes with expressive feature vectors such as molecular fingerprints and gene ontologies. Biologically and chemically relevant features improve the capacity of machine learning models to generate accurate and well-structured embedding spaces. The resulting resource represents a coherent and reliable biological knowledge graph that serves as a state-of-the-art platform to advance research in computational biology and precision medicine. Moreover, it offers the opportunity to benchmark graph-based machine learning and network medicine models on relevant tasks. We demonstrate the effectiveness of the proposed dataset by benchmarking it against the task of drug repurposing, PPI prediction, and side-effect prediction, modeled as link prediction problems.

知识图谱药物重定位生物信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。