首个专为空间转录组设计的图文对齐基础模型,融合病理图像与基因表达信息。
ST-Align: A Multimodal Foundation Model for Image-Gene Alignment in Spatial Transcriptomics
- 通过多尺度与跨层级对齐策略,融合斑点与微环境上下文信息。
- 在6个数据集上实现优异零样本与少样本性能,支持130万对图像-基因数据预训练。
- 适合生物医学研究者探索组织中基因与结构的关系,降低实验成本。
空间转录组学(ST)在全片尺度上提供高分辨率病理图像和单个位点的全转录组表达谱,是构建多模态基础模型的理想数据源。尽管近期研究尝试基于位点级数据微调视觉编码器并加入可训练基因编码器,但缺乏更广泛的切片视角及空间内在关系,限制了其捕捉ST特有洞察的能力。本文提出ST-Align,首个专为空间转录组设计的基础模型,通过引入空间上下文深度对齐图像与基因信息,有效连接病理成像与基因特征。我们设计了一种新颖的预训练框架,采用三重对齐策略:(1)跨多尺度对齐图像-基因对,捕获位点级与微环境级上下文,获得全面视角;(2)跨层级对齐多模态见解,关联局部细胞特征与整体组织架构。此外,ST-Align采用针对不同ST上下文定制的编码器,并结合注意力融合网络(ABFN)实现增强的多模态融合,有效整合领域共性知识与空间转录组特异性信息。我们在130万对位点-微环境组合上预训练该模型,并在六个数据集上评估其在两个下游任务中的表现,结果表明其具备卓越的零样本与少样本能力。ST-Align展示了降低空间转录组研究成本、揭示人类组织关键组成差异的巨大潜力。
原文摘要 · Abstract (English)
Spatial transcriptomics (ST) provides high-resolution pathological images and whole-transcriptomic expression profiles at individual spots across whole-slide scales. This setting makes it an ideal data source to develop multimodal foundation models. Although recent studies attempted to fine-tune visual encoders with trainable gene encoders based on spot-level, the absence of a wider slide perspective and spatial intrinsic relationships limits their ability to capture ST-specific insights effectively. Here, we introduce ST-Align, the first foundation model designed for ST that deeply aligns image-gene pairs by incorporating spatial context, effectively bridging pathological imaging with genomic features. We design a novel pretraining framework with a three-target alignment strategy for ST-Align, enabling (1) multi-scale alignment across image-gene pairs, capturing both spot- and niche-level contexts for a comprehensive perspective, and (2) cross-level alignment of multimodal insights, connecting localized cellular characteristics and broader tissue architecture. Additionally, ST-Align employs specialized encoders tailored to distinct ST contexts, followed by an Attention-Based Fusion Network (ABFN) for enhanced multimodal fusion, effectively merging domain-shared knowledge with ST-specific insights from both pathological and genomic data. We pre-trained ST-Align on 1.3 million spot-niche pairs and evaluated its performance through two downstream tasks across six datasets, demonstrating superior zero-shot and few-shot capabilities. ST-Align highlights the potential for reducing the cost of ST and providing valuable insights into the distinction of critical compositions within human tissue.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。