用知识图谱融合多源基因表达数据,提升疾病诊断准确率
Multi-dataset and Transfer Learning Using Gene Expression Knowledge Graphs
- 构建基因表达知识图谱,整合多数据集与领域知识
- 在三种学习场景下均实现诊断性能提升
- 适合生物医学数据融合与精准医疗研究者
基因表达数据可揭示基因调控机制、生化通路及细胞功能,通过对比疾病与健康患者的表达谱,有助于理解疾病病理。因此,机器学习被广泛用于处理基因表达数据,患者诊断成为最热门的应用之一。然而,基因表达数据通常样本量有限,且不同数据集间基因表达差异大,难以直接合并。本文提出一种新方法,通过知识图谱整合多个基因表达数据集与领域知识,利用图嵌入技术生成向量表示,输入图神经网络和多层感知机进行建模。在单数据集学习、多数据集学习和迁移学习三种设置下评估效果,结果表明,融合数据与知识显著提升了患者诊断性能。
原文摘要 · Abstract (English)
Gene expression datasets offer insights into gene regulation mechanisms, biochemical pathways, and cellular functions. Additionally, comparing gene expression profiles between disease and control patients can deepen the understanding of disease pathology. Therefore, machine learning has been used to process gene expression data, with patient diagnosis emerging as one of the most popular applications. Although gene expression data can provide valuable insights, challenges arise because the number of patients in expression datasets is usually limited, and the data from different datasets with different gene expressions cannot be easily combined. This work proposes a novel methodology to address these challenges by integrating multiple gene expression datasets and domain-specific knowledge using knowledge graphs, a unique tool for biomedical data integration. Then, vector representations are produced using knowledge graph embedding techniques, which are used as inputs for a graph neural network and a multi-layer perceptron. We evaluate the efficacy of our methodology in three settings: single-dataset learning, multi-dataset learning, and transfer learning. The experimental results show that combining gene expression datasets and domain-specific knowledge improves patient diagnosis in all three settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。