用空间转录组数据增强病理图像模型,提升分子预测与跨模态能力。
Towards Spatial Transcriptomics-driven Pathology Foundation Models
- 通过基因表达点与组织区域配对,自监督融合分子信息到病理图像编码器。
- 在38个切片级和15个斑块级任务中,性能超越纯视觉与转录组基线。
- 适配现有病理大模型,支持基因到图像检索等新功能,可迁移性强。
空间转录组(ST)提供基因表达的空间定位测量,使人类组织的分子景观表征超越传统组织学评估,并能与形态学对齐。多模态基础模型的成功表明,局部表达与形态之间的形态-分子耦合可系统性地改进组织学表示。我们提出空间表达对齐学习(SEAL),一种将局部分子信息融入病理视觉编码器的视觉-组学自监督学习框架。SEAL不从头训练编码器,而是作为参数高效的视觉-组学微调方法,可灵活应用于广泛使用的病理基础模型。我们在超过70万对基因表达点-组织区域样本上训练,涵盖14个器官的肿瘤与正常样本。在38个切片级和15个斑块级下游任务中测试,SEAL作为病理基础模型的即插即用替代方案,持续优于广泛使用的纯视觉和ST预测基线,在切片级分子状态、通路活性及治疗反应预测,以及斑块级基因表达预测任务中表现更优。此外,SEAL编码器在分布外评估中表现出强泛化能力,并实现基因到图像检索等新型跨模态功能。本工作提出了一种基于空间转录组引导的病理基础模型微调通用框架,证明引入局部分子监督是提升视觉表示并拓展其跨模态用途的有效且可行路径。
原文摘要 · Abstract (English)
Spatial transcriptomics (ST) provides spatially resolved measurements of gene expression, enabling characterization of the molecular landscape of human tissue beyond histological assessment as well as localized readouts that can be aligned with morphology. Concurrently, the success of multimodal foundation models that integrate vision with complementary modalities suggests that morphomolecular coupling between local expression and morphology can be systematically used to improve histological representations themselves. We introduce Spatial Expression-Aligned Learning (SEAL), a vision-omics self-supervised learning framework that infuses localized molecular information into pathology vision encoders. Rather than training new encoders from scratch, SEAL is designed as a parameter-efficient vision-omics finetuning method that can be flexibly applied to widely used pathology foundation models. We instantiate SEAL by training on over 700,000 paired gene expression spot-tissue region examples spanning tumor and normal samples from 14 organs. Tested across 38 slide-level and 15 patch-level downstream tasks, SEAL provides a drop-in replacement for pathology foundation models that consistently improves performance over widely used vision-only and ST prediction baselines on slide-level molecular status, pathway activity, and treatment response prediction, as well as patch-level gene expression prediction tasks. Additionally, SEAL encoders exhibit robust domain generalization on out-of-distribution evaluations and enable new cross-modal capabilities such as gene-to-image retrieval. Our work proposes a general framework for ST-guided finetuning of pathology foundation models, showing that augmenting existing models with localized molecular supervision is an effective and practical step for improving visual representations and expanding their cross-modal utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。