将文献知识动态融入图神经网络,提升癌症信号通路功能聚类效果。
RAG-GNN: Integrating Retrieved Knowledge with Graph Neural Networks for Precision Medicine
- 用检索投影与门控融合机制,实时结合文献知识增强图网络表征。
- 功能聚类轮廓系数提升0.093,检索准确率@10达0.242(提升152%)。
- 适合精准医疗中需融合生物文献与网络结构的研究者使用。
网络拓扑擅长结构预测,但难以捕捉生物文献中的功能语义。我们提出RAG-GNN,一种端到端可训练的检索增强图神经网络框架,通过联合优化的检索投影、门控融合机制和对比对齐,将GNN表示与动态检索的文献知识相融合。在包含379个蛋白、3,498个相互作用、14个功能类别的一次癌症信号通路案例研究中,RAG-GNN将功能聚类轮廓系数从GNN-only的-0.237 ± 0.065提升至-0.144 ± 0.066,平均提升0.093 ± 0.022(10次随机种子)。所学检索方法的均值精确率@10为0.242,较随机基线(0.096)提升152%。通过自助法置信区间分析发现,拓扑与检索编码了95.6%的共享信息,且检索同时提升了簇内凝聚力(轮廓系数)与簇一致性(ARI +0.021 ± 0.015)。反事实实验表明,对抗性、缺失或随机检索均导致性能下降,验证了门控融合机制依赖于文档内容。与八种现有嵌入方法对比显示任务特异性互补:拓扑导向方法在链接预测表现优异,而检索增强则持续改善功能聚类。DDR1子网络分析结果与已知合成致死关系一致。这些结果表明,拓扑仅与检索增强方法在精准医疗中具有互补价值。
原文摘要 · Abstract (English)
Network topology excels at structural predictions but fails to capture functional semantics encoded in biomedical literature. We present RAG-GNN, an end-to-end trainable retrieval-augmented graph neural network framework that integrates GNN representations with dynamically retrieved literature-derived knowledge through a jointly optimized retrieval projection, gated fusion mechanism, and contrastive alignment. In a cancer signaling case study (379 proteins, 3,498 interactions, 14 functional categories), RAG-GNN improves functional clustering from silhouette $= -0.237 \pm 0.065$ (GNN-only) to $-0.144 \pm 0.066$, a consistent improvement of $+0.093 \pm 0.022$ across 10 random seeds, while the learned retrieval achieves mean precision@10 $= 0.242$, a 152\% improvement over the random baseline ($0.096$). Heuristic information decomposition with bootstrap confidence intervals reveals that topology and retrieval encode overwhelmingly shared information (95.6\%), with retrieval improving both intra-cluster cohesion (silhouette) and cluster agreement (ARI $+0.021 \pm 0.015$). Counterfactual experiments confirm that adversarial, absent, and random retrieval all degrade performance, validating that the gated fusion mechanism depends on document content. Benchmarking against eight established embedding methods demonstrates task-specific complementarity: topology-focused methods achieve strong link prediction, while retrieval augmentation consistently improves functional clustering within the controlled GNN-only ablation. DDR1 subnetwork analysis provides confirmatory validation consistent with established synthetic lethality relationships. These results establish that topology-only and retrieval-augmented approaches serve complementary purposes for precision medicine applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。