arXiv:2601.15504cs.LGq-bio.GN2026-01

轻量图网络模型精准还原空间基因表达,助力病理分析与调控推断。

SAGE-FM: A lightweight and interpretable spatial transcriptomics foundation model

  • 基于图卷积网络,通过预测被遮蔽区域基因表达建模空间关系。
  • 91%被遮蔽基因恢复后相关性显著(p<0.05),优于现有方法。
  • 适合生物医学研究者用于空间转录组下游任务与机制探索。

空间转录组学可实现空间定位的基因表达分析,推动了对空间依赖调控关系建模的计算方法发展。本文提出SAGE-FM,一种基于图卷积网络(GCN)的轻量级空间转录组基础模型,采用遮蔽中心点预测目标进行训练。模型在涵盖15个器官的416个人类Visium样本上训练,学习到具有空间一致性的嵌入表示,能稳健恢复被遮蔽基因,其中91%的被遮蔽基因在统计上表现出显著相关性(p < 0.05)。SAGE-FM生成的嵌入在无监督聚类和生物异质性保留方面优于MOFA及现有空间转录组方法。该模型在下游任务中具备泛化能力:在口咽鳞状细胞癌中实现81%的病理学家定义点注释准确率,并改进胶质母细胞瘤亚型预测性能。体外扰动实验进一步验证模型捕捉到了与真实情况一致的配体-受体方向性及上下游调控效应。结果表明,简单且参数高效的GCN可作为大规模空间转录组的生物学可解释、空间感知的基础模型。

原文摘要 · Abstract (English)

Spatial transcriptomics enables spatial gene expression profiling, motivating computational models that capture spatially conditioned regulatory relationships. We introduce SAGE-FM, a lightweight spatial transcriptomics foundation model based on graph convolutional networks (GCNs) trained with a masked central spot prediction objective. Trained on 416 human Visium samples spanning 15 organs, SAGE-FM learns spatially coherent embeddings that robustly recover masked genes, with 91% of masked genes showing significant correlations (p < 0.05). The embeddings generated by SAGE-FM outperform MOFA and existing spatial transcriptomics methods in unsupervised clustering and preservation of biological heterogeneity. SAGE-FM generalizes to downstream tasks, enabling 81% accuracy in pathologist-defined spot annotation in oropharyngeal squamous cell carcinoma and improving glioblastoma subtype prediction relative to MOFA. In silico perturbation experiments further demonstrate that the model captures directional ligand-receptor and upstream-downstream regulatory effects consistent with ground truth. These results demonstrate that simple, parameter-efficient GCNs can serve as biologically interpretable and spatially aware foundation models for large-scale spatial transcriptomics.

空间转录组图神经网络基础模型生物可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。