用基因表达信息增强病理图像学习,提升模型对空间结构的感知能力。
Fusing Pixels and Genes: Spatially-Aware Learning in Computational Pathology
- 引入空间基因数据,实现图像与转录组的联合嵌入。
- 在六个数据集上跨任务表现优异,显著提升模型泛化性。
- 适合从事病理多模态研究、空间转录组分析的研究者。
近年来,计算病理学中的多模态学习取得显著进展,但现有模型主要依赖视觉与语言模态,语言缺乏分子特异性且病理监督有限,导致表征瓶颈。本文提出STAMP框架,通过整合空间解析基因表达谱,实现分子引导的病理图像与转录组联合表示学习。自监督基因引导训练提供鲁棒、任务无关的信号,结合空间上下文与多尺度信息,显著提升模型性能与泛化能力。为此,我们构建了目前最大的基于Visium的空间转录组数据集SpaVis-6M,并在此基础上训练空间感知基因编码器。通过分层多尺度对比对齐与跨尺度块定位机制,STAMP有效对齐空间转录组与病理图像,捕捉空间结构与分子变异。我们在六个数据集和四个下游任务中验证了STAMP,结果一致表现优异。这凸显了空间解析分子监督在推动计算病理多模态学习中的价值与必要性。代码见附录,预训练权重与SpaVis-6M可于https://github.com/Hanminghao/STAMP获取。
原文摘要 · Abstract (English)
Recent years have witnessed remarkable progress in multimodal learning within computational pathology. Existing models primarily rely on vision and language modalities; however, language alone lacks molecular specificity and offers limited pathological supervision, leading to representational bottlenecks. In this paper, we propose STAMP, a Spatial Transcriptomics-Augmented Multimodal Pathology representation learning framework that integrates spatially-resolved gene expression profiles to enable molecule-guided joint embedding of pathology images and transcriptomic data. Our study shows that self-supervised, gene-guided training provides a robust and task-agnostic signal for learning pathology image representations. Incorporating spatial context and multi-scale information further enhances model performance and generalizability. To support this, we constructed SpaVis-6M, the largest Visium-based spatial transcriptomics dataset to date, and trained a spatially-aware gene encoder on this resource. Leveraging hierarchical multi-scale contrastive alignment and cross-scale patch localization mechanisms, STAMP effectively aligns spatial transcriptomics with pathology images, capturing spatial structure and molecular variation. We validate STAMP across six datasets and four downstream tasks, where it consistently achieves strong performance. These results highlight the value and necessity of integrating spatially resolved molecular supervision for advancing multimodal learning in computational pathology. The code is included in the supplementary materials. The pretrained weights and SpaVis-6M are available at: https://github.com/Hanminghao/STAMP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。