用生物知识图谱增强临床数据模型,小样本下预测更准。
Knowledge Graph Modulated Deep Learning for Limited-Sample Clinical Data Analysis

- 将患者建模为模块化图,融合基因表达与通路知识图谱
- 在近万人数据上,前列腺癌诊断宏F1提升49个百分点
- 适合小样本临床研究,且结果可解释性强
生物系统由结构化的分子互作支配,通路、调控回路和基因功能关系决定细胞行为与疾病进展。这些知识天然以图的形式存在。然而,大多数生物医学AI模型无法直接利用图编码的生物学知识,需将其压缩为低维表示,易丢失关键结构并降低性能,尤其在小样本临床研究中。本文提出图中图(GiG)框架,一种知识图谱调制的深度学习方法,用于数据高效临床预测。GiG将每位患者表示为独立模块化图,其中人工梳理的生物知识图谱定义边,患者特异性测量如基因表达定义节点特征。该设计支持多知识图谱融合,同时保留基因-基因互作与通路拓扑结构。在包含近9,700名患者的多个队列及五项临床任务中,包括液体活检癌症检测、前列腺癌诊断和32类泛癌分类,GiG持续优于传统与先进方法,尤其在小样本设置下表现突出。在挑战性的前列腺癌诊断任务中,相比竞品方法,宏观F1最高提升49个百分点。控制实验显示,替换真实通路图为随机拓扑后性能下降,证实收益源于生物学基础的图结构而非图建模本身。结果表明,知识图谱调制的深度学习能提升临床数据分析的鲁棒性、可解释性与样本效率,并为整合生物知识图谱提供原则性框架。
原文摘要 · Abstract (English)
Biological systems are governed by structured molecular interactions, where pathways, regulatory circuits, and functional gene relationships shape cellular behavior and disease progression. Much of this knowledge is naturally represented as graphs. However, most biomedical AI models cannot directly use graph-encoded biological knowledge and instead require compressed low-dimensional representations, which can lose important structure and reduce performance, especially in limited-sample clinical studies. Here, we introduce Graph-in-Graph (GiG), a knowledge graph-modulated deep learning framework for data-efficient clinical prediction. GiG represents each patient as a standalone modular graph, in which curated biological knowledge graphs define edges and patient-specific measurements, such as gene expression, define node features. This design allows multiple biological knowledge graphs to be integrated while preserving gene-gene interactions and pathway topology during patient-level representation learning. Across cohorts comprising nearly 9,700 patients and five clinical tasks, including liquid biopsy cancer detection, prostate cancer diagnosis, and 32-class pan-cancer classification, GiG consistently outperforms traditional and state-of-the-art methods, with the largest gains in limited-sample settings. On the challenging prostate cancer diagnosis task, GiG improves macro-F1 by up to 49 percentage points relative to competing methods. Control experiments replacing real pathway graphs with random topologies confirm that these gains arise from biologically grounded knowledge graph structure rather than graph modeling alone. These findings show that knowledge graph-modulated deep learning can improve robustness, interpretability, and sample efficiency in clinical data analysis, and provide a principled framework for integrating biological knowledge graphs into predictive modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。