用基因敲降实验数据做监督,让图对比学习更贴近真实生物机制。
Supervised Graph Contrastive Learning for Gene Regulatory Networks
- 将真实基因敲降实验作为监督信号,替代传统人工扰动
- 在三种癌症患者来源的网络上,疾病亚型区分更清晰,聚类效果更好
- 适合研究基因调控网络、癌症分型等需要生物学可解释性的任务
图对比学习(GCL)是一种强大的自监督学习框架,通过图扰动进行数据增强,广泛应用于基因调控网络(GRNs)分析。但现有GCL常用的人工扰动(如节点删除)会引入与生物现实不符的结构变化,促使图表示学习向无增强方法发展。然而,这种趋势忽视了一个关键洞察:来自生物有意义扰动的结构变化并非问题,反而是信息富集的来源,忽略了利用真实生物实验数据的潜力。为此,我们提出SupGCL(监督图对比学习),一种针对GRNs的新GCL方法,直接将基因敲降实验中的生物扰动作为监督信号。SupGCL采用概率建模,连续推广传统GCL,将人工增强与真实敲降实验测量结果相联系,并以真实数据作为显式监督。在三种癌症类型的患者来源GRN上,使用SupGCL训练网络表示,在两个评估场景中表现优异:(i) 嵌入空间分析中,揭示更清晰的疾病亚型结构并提升聚类性能;(ii) 任务特定微调中,在13项下游任务(涵盖基因层面功能注释与患者层面预测)上持续优于多个强基线模型。
原文摘要 · Abstract (English)
Graph Contrastive Learning (GCL) is a powerful self-supervised learning framework that performs data augmentation through graph perturbations, with growing applications in the analysis of biological networks such as Gene Regulatory Networks (GRNs). The artificial perturbations commonly used in GCL, such as node dropping, induce structural changes that can diverge from biological reality. This concern has contributed to a broader trend in graph representation learning toward augmentation-free methods, which view such structural changes as problematic and should be avoided. However, this trend overlooks the fundamental insight that structural changes from biologically meaningful perturbations are not a problem to be avoided, but rather a rich source of information, thereby ignoring the valuable opportunity to leverage data from real biological experiments. Motivated by this insight, we propose SupGCL (Supervised Graph Contrastive Learning), a new GCL method for GRNs that directly incorporates biological perturbations from gene knockdown experiments as supervision. SupGCL is a probabilistic formulation that continuously generalizes conventional GCL, linking artificial augmentations with real perturbations measured in knockdown experiments, and using the latter as explicit supervision. On patient-derived GRNs from three cancer types, we train GRN representations with SupGCL and evaluate it in two regimes: (i) embedding space analysis, where it yields clearer disease-subtype structure and improves clustering, and (ii) task-specific fine-tuning, where it consistently outperforms strong graph representation learning baselines on 13 downstream tasks spanning gene-level functional annotation and patient-level prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。