arXiv:2508.01490q-bio.GNcs.AI2025-08ICCV被引 10

构建首个空间转录组跨模态学习大基准,揭示预训练策略的矛盾效果

A Large-Scale Benchmark of Cross-Modal Learning for Histology and Gene Expression in Spatial Transcriptomics

  • 基于6种基因面板、54个捐献者数据构建跨模态对比预训练基准
  • 基因表达模型性能主导表征对齐,但对比预训练反而降低基因表达预测精度
  • 发现批次效应是干扰对齐的关键因素,适合多模态算法研究者使用

空间转录组学可同时测量基因表达与组织形态,为细胞结构和疾病机制提供新视角。然而,该领域缺乏评估多模态学习方法的综合性基准。本文提出HESCAPE,一个基于6种基因面板、54名捐献者的泛器官数据集,用于空间转录组中跨模态对比预训练的大规模基准。系统评估了当前主流图像与基因表达编码器在多种预训练策略下的表现,并在基因突变分类与基因表达预测两个下游任务上进行测试。结果表明,基因表达编码器是决定表征对齐能力的关键因素;在空间转录组数据上预训练的基因模型优于无空间数据训练或简单基线模型。然而,下游任务评估显示显著矛盾:尽管对比预训练持续提升基因突变分类性能,却导致基因表达预测性能低于未采用跨模态目标的基线编码器。我们识别出批次效应是干扰有效跨模态对齐的主要因素。研究强调空间转录组学中需发展抗批次的多模态学习方法。为推动该方向进展,我们发布HESCAPE,提供标准化数据集、评估协议与基准工具。

原文摘要 · Abstract (English)

Spatial transcriptomics enables simultaneous measurement of gene expression and tissue morphology, offering unprecedented insights into cellular organization and disease mechanisms. However, the field lacks comprehensive benchmarks for evaluating multimodal learning methods that leverage both histology images and gene expression data. Here, we present HESCAPE, a large-scale benchmark for cross-modal contrastive pretraining in spatial transcriptomics, built on a curated pan-organ dataset spanning 6 different gene panels and 54 donors. We systematically evaluated state-of-the-art image and gene expression encoders across multiple pretraining strategies and assessed their effectiveness on two downstream tasks: gene mutation classification and gene expression prediction. Our benchmark demonstrates that gene expression encoders are the primary determinant of strong representational alignment, and that gene models pretrained on spatial transcriptomics data outperform both those trained without spatial data and simple baseline approaches. However, downstream task evaluation reveals a striking contradiction: while contrastive pretraining consistently improves gene mutation classification performance, it degrades direct gene expression prediction compared to baseline encoders trained without cross-modal objectives. We identify batch effects as a key factor that interferes with effective cross-modal alignment. Our findings highlight the critical need for batch-robust multimodal learning approaches in spatial transcriptomics. To accelerate progress in this direction, we release HESCAPE, providing standardized datasets, evaluation protocols, and benchmarking tools for the community

空间转录组跨模态学习基准测试基因表达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。