arXiv:2605.21502q-bio.MNcs.AI2026-05

通过融合图神经网络解释,发现癌症关键基因具有特定拓扑分布模式。

Graph neural network explanations reveal a topological signature of disease-associated hubs in biological networks

  • 结合壳层得分与多解释器共识,提升关键基因识别能力。
  • IG和LRP在局部中心节点上表现更好,能捕捉疾病枢纽的拓扑特征。
  • 该方法适合研究癌症通路和分子机制,尤其对低度节点有效。

图神经网络(GNN)在建模生物系统中日益普及,但后处理解释方法是否能可靠恢复有意义的分子机制仍不明确。本文系统评估了四种常用方法:敏感性归因(SA)、积分梯度(IG)、GNNExplainer 和逐层相关性传播(LRP),用于在乳腺癌RNA-seq数据投影到蛋白质互作网络中识别疾病相关结构。基于已知真值的合成基准测试显示,不同方法恢复信号组织方式各异:SA在稀疏单节点驱动情形下最优,而IG和LRP更擅长恢复分布式通路样与级联样信号。在TCGA BRCA数据中,我们发现疾病相关枢纽存在一致的拓扑签名:归因峰值集中在1跳邻域,并随网络层级衰减,这一模式在IG和LRP中尤为显著,且与已知癌症枢纽高度富集相关。此外,局部枢纽富集与全局基因排序性能之间存在权衡:IG优化局部富集,而SA实现更优全局区分。受此互补行为启发,我们提出一种框架,结合壳层式枢纽评分与多解释器共识排名。共识得分提升了经典癌基因(TP53、BRCA1、ESR1、MYC)的优先级,降低对节点度的依赖,尤其在调参后优于单一方法。通路富集分析进一步表明,该方法更准确恢复了包括ERBB2、RTK、MAPK、免疫及细胞因子信号在内的生物共性癌症程序。

原文摘要 · Abstract (English)

Graph neural networks (GNNs) are increasingly used to model biological systems, yet the reliability of post-hoc explanation methods for recovering meaningful molecular mechanisms remains unclear. Here, we systematically evaluate four widely used approaches: Saliency Attribution (SA), Integrated Gradients (IG), GNNExplainer, and Layer-wise Relevance Propagation (LRP) for identifying disease-relevant structure in breast cancer RNA-seq data projected onto a protein-protein interaction network. Using synthetic benchmarks with known ground-truth motifs, we show that explanation methods recover distinct signal organizations: SA performs best for sparse single-node drivers, whereas IG and LRP preferentially recover distributed pathway-like and cascade-like signals. In TCGA BRCA data, we identify a consistent topological signature of disease-associated hubs in which attribution peaks in the immediate 1-hop neighborhood and decays across successive network shells, a pattern most pronounced for IG and LRP and associated with strong enrichment of known cancer hubs. We further observe a trade-off between local hub enrichment and global gene ranking performance, with IG optimizing local enrichment and SA achieving superior global discrimination. Motivated by these complementary behaviors, we introduce a framework combining a shell-based hub score with consensus ranking across explainers. Consensus scores improve prioritization of canonical cancer genes (TP53, BRCA1, ESR1, MYC), reduce dependence on node degree, and, especially when tuned, outperform individual methods. Pathway enrichment further reveals improved recovery of biologically coherent cancer programs, including ERBB2, RTK, MAPK, immune, and cytokine signaling. Together, these results demonstrate that topology-aware integration of graph explanations can improve biological interpretability and biologically relevant molecular recovery.

图神经网络癌症机制生物网络解释方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。