构建多尺度空间转录组基础模型,提升组织分析精度
SToFM: a Multi-scale Foundation Model for Spatial Transcriptomics
- 通过多尺度信息提取构建融合宏观、微观与基因层的子切片
- 在组织区域分割和细胞类型注释任务上表现优异
- 适用于生物组织分析、疾病研究等需要精细解析的场景
空间转录组技术通过保留细胞的空间位置信息,为单细胞生物学研究提供了丰富洞见。构建空间转录组基础模型可显著提升对海量复杂数据的分析能力,揭示组织结构的细微特征。然而,建模空间转录组数据面临挑战,需从包含大量细胞的组织切片中提取多层次信息,整合宏观组织形态、微观细胞微环境及基因表达谱。为此,我们提出SToFM,一种多尺度空间转录组基础模型。该模型首先对每张空间转录组切片进行多尺度信息提取,生成融合宏观、微观与基因尺度信息的子切片;随后利用SE(2) Transformer从子切片中获取高质量细胞表示。此外,我们构建了目前最大规模的高分辨率空间转录组语料库SToCorpus-88M用于预训练。SToFM在多种下游任务(如组织区域语义分割、细胞类型注释)中表现卓越,证明其通过捕捉并整合多尺度信息,具备对空间转录组数据的全面理解能力。
原文摘要 · Abstract (English)
Spatial Transcriptomics (ST) technologies provide biologists with rich insights into single-cell biology by preserving spatial context of cells. Building foundational models for ST can significantly enhance the analysis of vast and complex data sources, unlocking new perspectives on the intricacies of biological tissues. However, modeling ST data is inherently challenging due to the need to extract multi-scale information from tissue slices containing vast numbers of cells. This process requires integrating macro-scale tissue morphology, micro-scale cellular microenvironment, and gene-scale gene expression profile. To address this challenge, we propose SToFM, a multi-scale Spatial Transcriptomics Foundation Model. SToFM first performs multi-scale information extraction on each ST slice, to construct a set of ST sub-slices that aggregate macro-, micro- and gene-scale information. Then an SE(2) Transformer is used to obtain high-quality cell representations from the sub-slices. Additionally, we construct \textbf{SToCorpus-88M}, the largest high-resolution spatial transcriptomics corpus for pretraining. SToFM achieves outstanding performance on a variety of downstream tasks, such as tissue region semantic segmentation and cell type annotation, demonstrating its comprehensive understanding of ST data through capturing and integrating multi-scale information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。