提出首个通用图结构基础模型,可跨领域迁移用于生物医学图数据
Toward a universal foundation model for graph-structured data
- 用度统计、中心性等无特征图属性生成结构提示,融合消息传递网络
- 在SagePPI上零样本/少样本性能超越基线21.8%,均值ROC-AUC达95.5%
- 一次预训练即可适配新数据集,适合跨机构、多模态图分析场景
图是生物医学研究的核心表示形式,用于刻画分子互作网络、基因调控回路、细胞间通讯图谱和知识图谱。尽管重要,目前尚无类似语言与视觉领域的通用图分析基础模型。现有图神经网络通常仅在单一数据集上训练,学习到的表征局限于特定节点特征、拓扑结构和标签空间,难以跨域迁移。这一局限在生物医学中尤为严重,因不同队列、检测方法和机构间的网络差异显著。本文提出一种图基础模型,旨在学习不依赖具体节点身份或特征方案的可迁移结构表征。该方法利用度统计、中心性度量、社区结构指标及基于扩散的签名等无特征图属性,构建结构提示,并将其与消息传递主干网络结合,将异构图嵌入共享表示空间。模型在异构图上一次性预训练后,可在未见数据集上以极小适应实现重用。在多个基准测试中,其性能匹配甚至超过强监督基线,且在未见图上展现出更优的零样本与少样本泛化能力。在SagePPI基准上,对预训练主干进行微调后,平均ROC-AUC达95.5%,较最优监督消息传递基线提升21.8%。该技术为生物医学与网络科学中的可复用、基础规模图模型提供了全新路径。
原文摘要 · Abstract (English)
Graphs are a central representation in biomedical research, capturing molecular interaction networks, gene regulatory circuits, cell--cell communication maps, and knowledge graphs. Despite their importance, currently there is not a broadly reusable foundation model available for graph analysis comparable to those that have transformed language and vision. Existing graph neural networks are typically trained on a single dataset and learn representations specific only to that graph's node features, topology, and label space, limiting their ability to transfer across domains. This lack of generalization is particularly problematic in biology and medicine, where networks vary substantially across cohorts, assays, and institutions. Here we introduce a graph foundation model designed to learn transferable structural representations that are not specific to specific node identities or feature schemes. Our approach leverages feature-agnostic graph properties, including degree statistics, centrality measures, community structure indicators, and diffusion-based signatures, and encodes them as structural prompts. These prompts are integrated with a message-passing backbone to embed diverse graphs into a shared representation space. The model is pretrained once on heterogeneous graphs and subsequently reused on unseen datasets with minimal adaptation. Across multiple benchmarks, our pretrained model matches or exceeds strong supervised baselines while demonstrating superior zero-shot and few-shot generalization on held-out graphs. On the SagePPI benchmark, supervised fine-tuning of the pretrained backbone achieves a mean ROC-AUC of 95.5%, a gain of 21.8% over the best supervised message-passing baseline. The proposed technique thus provides a unique approach toward reusable, foundation-scale models for graph-structured data in biomedical and network science applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。