arXiv:2505.11023cs.LG2025-05

在生物医学领域,背景知识对图神经网络性能提升有限,甚至无效。

Informed, but Not Always Improved: Challenging the Benefit of Background Knowledge in GNNs

  • 用合成数据和扰动实验检验背景知识对GNN的影响
  • 真实数据中引入背景知识后模型表现未提升,扰动也无显著影响
  • 需匹配图神经网络架构与背景知识特性才能发挥潜力

在复杂且数据稀少的生物医学研究领域,将背景知识图(如蛋白质互作网络)融入基于图的机器学习流程是一项有前景的方向。然而,尽管普遍认为背景知识能提升模型性能,其实际贡献及不完美知识的影响仍不明确。本文研究了背景知识在癌症亚型分类这一重要现实任务中的作用。令人惊讶的是,我们发现:(i) 使用背景知识的先进GNN模型表现不优于线性回归等无知识模型;(ii) 即使背景知识图被严重扰动,模型性能也基本不变。为理解这些反直觉结果,我们提出一个评估框架,包含(i)背景知识明确有效的合成场景,以及(ii)模拟背景知识图各种缺陷的扰动集合。通过该框架,在合成与真实生物医学场景中测试了背景知识感知模型的鲁棒性。结果表明,必须精细匹配图神经网络架构与背景知识特征,才能实现显著性能提升。

原文摘要 · Abstract (English)

In complex and low-data domains such as biomedical research, incorporating background knowledge (BK) graphs, such as protein-protein interaction (PPI) networks, into graph-based machine learning pipelines is a promising research direction. However, while BK is often assumed to improve model performance, its actual contribution and the impact of imperfect knowledge remain poorly understood. In this work, we investigate the role of BK in an important real-world task: cancer subtype classification. Surprisingly, we find that (i) state-of-the-art GNNs using BK perform no better than uninformed models like linear regression, and (ii) their performance remains largely unchanged even when the BK graph is heavily perturbed. To understand these unexpected results, we introduce an evaluation framework, which employs (i) a synthetic setting where the BK is clearly informative and (ii) a set of perturbations that simulate various imperfections in BK graphs. With this, we test the robustness of BK-aware models in both synthetic and real-world biomedical settings. Our findings reveal that careful alignment of GNN architectures and BK characteristics is necessary but holds the potential for significant performance improvements.

图神经网络背景知识生物医学鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。