提出多解释高阶知识图谱模型,提升癌症合成致死预测可信度
Interpretable High-order Knowledge Graph Neural Network for Predicting Synthetic Lethality in Human Cancers
- 设计新信息瓶颈机制,结合点过程约束生成多个忠实解释
- 使用13种基序邻接矩阵,有效编码基因关系的高阶结构
- 相比现有方法提升性能,揭示多种潜在生物机制
合成致死(SL)是癌症治疗中的重要基因互作机制。现有方法将知识图谱融入图神经网络,用注意力机制提取局部子图作为解释,但存在解释不可靠、仅生成单一解释、难以保证高阶结构可信等问题。为此,我们提出DGIB4SL——一种基于知识图谱的GNN模型,能为同一基因对生成多个忠实解释,并有效编码高阶结构。核心在于引入新型DGIB目标函数,将行列式点过程(DPP)约束融入标准信息瓶颈框架,并采用13种基序相关的邻接矩阵捕捉基因表征中的高阶结构。实验表明,DGIB4SL超越当前最优基线,在多个数据集上表现更优,且能提供多角度解释,揭示了合成致死推断背后的多样生物学机制。
原文摘要 · Abstract (English)
Synthetic lethality (SL) is a promising gene interaction for cancer therapy. Recent SL prediction methods integrate knowledge graphs (KGs) into graph neural networks (GNNs) and employ attention mechanisms to extract local subgraphs as explanations for target gene pairs. However, attention mechanisms often lack fidelity, typically generate a single explanation per gene pair, and fail to ensure trustworthy high-order structures in their explanations. To overcome these limitations, we propose Diverse Graph Information Bottleneck for Synthetic Lethality (DGIB4SL), a KG-based GNN that generates multiple faithful explanations for the same gene pair and effectively encodes high-order structures. Specifically, we introduce a novel DGIB objective, integrating a Determinant Point Process (DPP) constraint into the standard IB objective, and employ 13 motif-based adjacency matrices to capture high-order structures in gene representations. Experimental results show that DGIB4SL outperforms state-of-the-art baselines and provides multiple explanations for SL prediction, revealing diverse biological mechanisms underlying SL inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。