对比两种方法预测基因与疾病关联,发现链接预测更优。
A Systematic Evaluation of Knowledge Graph Embeddings for Gene-Disease Association Prediction
- 用知识图谱嵌入进行链接预测或节点对分类
- 补充疾病本体和跨本体链接显著提升效果
- 链接预测法整体表现更好,适合生物医学关联挖掘
发现基因-疾病关联在生物学和医学中至关重要,有助于疾病识别与药物重定位。机器学习方法通过利用本体和知识图谱的结构加速这一过程。然而,许多现有工作忽略了显式表示疾病的本体,遗漏了它们之间的因果与语义关系。基因-疾病关联问题可自然建模为链接预测任务,嵌入算法通过探索知识图谱的结构和属性直接预测关联;也有研究将其视为节点对分类任务,结合嵌入与传统机器学习算法。该策略符合机器学习流程逻辑,但负样本使用及缺乏验证的基因-疾病关联限制了其效果。本文提出一个系统性框架,比较链接预测与节点对分类任务的表现,分析前沿基因-疾病关联方法性能,并评估不同顺序形式化的影响。还考察了通过疾病特异性本体提升语义丰富性及本体间额外链接的作用。框架包含五个步骤:数据划分、知识图谱整合、嵌入、建模与预测、方法评估。结果表明,增强疾病语义表示略有提升,而额外链接影响更大。链接预测方法能更好挖掘知识图谱中的语义信息。尽管节点对分类方法识别出所有真实阳性,链接预测方法总体表现更优。
原文摘要 · Abstract (English)
Discovery gene-disease links is important in biology and medicine areas, enabling disease identification and drug repurposing. Machine learning approaches accelerate this process by leveraging biological knowledge represented in ontologies and the structure of knowledge graphs. Still, many existing works overlook ontologies explicitly representing diseases, missing causal and semantic relationships between them. The gene-disease association problem naturally frames itself as a link prediction task, where embedding algorithms directly predict associations by exploring the structure and properties of the knowledge graph. Some works frame it as a node-pair classification task, combining embedding algorithms with traditional machine learning algorithms. This strategy aligns with the logic of a machine learning pipeline. However, the use of negative examples and the lack of validated gene-disease associations to train embedding models may constrain its effectiveness. This work introduces a novel framework for comparing the performance of link prediction versus node-pair classification tasks, analyses the performance of state of the art gene-disease association approaches, and compares the different order-based formalizations of gene-disease association prediction. It also evaluates the impact of the semantic richness through a disease-specific ontology and additional links between ontologies. The framework involves five steps: data splitting, knowledge graph integration, embedding, modeling and prediction, and method evaluation. Results show that enriching the semantic representation of diseases slightly improves performance, while additional links generate a greater impact. Link prediction methods better explore the semantic richness encoded in knowledge graphs. Although node-pair classification methods identify all true positives, link prediction methods outperform overall.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。