首个联合图文基因表达的预训练模型,提升空间转录组分析精度。
Learning from Gene Names, Expression Values and Images: Contrastive Masked Text-Image Pretraining for Spatial Transcriptomics Representation Learning
- 三模态联合建模:图像、基因名与表达值同步学习。
- 在多个数据集上实现超越基线的基因预测性能,支持零样本预测。
- 适合做空间转录组下游任务的研究者,尤其是跨实验泛化场景。
空间转录组学旨在将高分辨率组织学图像与空间分辨的基因表达数据关联。为提升基因表达预测等下游任务表现,需大规模预训练以获得跨组织、实验流程和实验室的通用表征。现有方法仅依赖基因名或表达值之一,削弱了基因语义并破坏基因与其数值的对应关系;同时,仅通过图像-文本对齐限制了视觉线索的利用。本文提出 CoMTIP,首个联合图像、基因名与表达值的对比掩码图文预训练框架,通过掩码特征建模重建遮蔽图像块,学习上下文感知的图像嵌入;文本分支采用可扩展的基因-文本编码器,平行处理所有基因句子,为每个基因及其数值分配专属嵌入,并使用成对对抗训练(PAAT)保持正确基因-数值关联。图像与文本表示在共享的 InfoNCE 优化空间中对齐。在公开空间转录组数据集上的实验表明,CoMTIP 不仅在多种下游任务中优于现有方法,还实现了零样本基因表达预测,这是此前方法不具备的能力。
原文摘要 · Abstract (English)
Spatial transcriptomics aims to connect high-resolution histology images with spatially resolved gene expression. To achieve better performance on downstream tasks such as gene expression prediction, large-scale pre-training is required to obtain generalisable representations that can bridge histology and transcriptomics across tissues, protocols, and laboratories. Existing cross-modal pre-training approaches for spatial transcriptomics rely on either gene names or expression values in isolation, which strips the gene branch of essential semantics and breaks the association between each gene and its quantitative magnitude. In addition, by restricting supervision to image-text alignment, these methods ignore intrinsic visual cues that are critical for learning robust image features. We present CoMTIP, the first Contrastive Masked Text-Image Pretraining framework that jointly learns from images, gene names, and expression values while capturing fine-grained visual context for spatial transcriptomics. The vision branch uses Masked Feature Modeling to reconstruct occluded patches and learn context-aware image embeddings. The text branch applies a scalable Gene-Text Encoder that processes all gene sentences in parallel, enriches each gene and its numerical value with dedicated embeddings, and employs Pair-aware Adversarial Training (PAAT) to preserve correct gene-value associations. Image and text representations are aligned in a shared InfoNCE-optimised space. Experiments on public spatial transcriptomics datasets show that CoMTIP not only surpasses previous methods on diverse downstream tasks but also achieves zero-shot gene expression prediction, a capability that existing approaches do not provide.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。