arXiv:2606.21156cs.CVcs.AI2026-06

用少量基因数据指导全片基因表达预测,提升病理图像分析精度。

Contrastive and Adaptive Multi-modal Masked Autoencoder for Spatial Transcriptomics

论文配图:Contrastive and Adaptive Multi-modal Masked Autoencoder for Spatial Transcriptomics
图 1 · 摘自论文原文
  • 将基因预测视为空间补全问题,用掩码自编码器结合视觉与基因模态
  • 仅需10%基因覆盖率即可达到最优效果,且无需基因数据时也更准确
  • 自适应选择关键组织点,适合真实病理设备使用场景

空间转录组学(ST)成本高昂,促使研究者尝试从H&E染色图像直接预测基因表达。然而,仅靠组织形态信息难以完全还原基因表达。为此,我们提出一种对比与自适应的多模态掩码自编码器(CAMMST),将该任务视为空间补全问题,利用少量基因表达作为遗传锚点来推断整张切片的基因表达谱。通过生物显著性评分和学习排序策略,自适应识别组织中最具信息量的区域,并选取连续区域作为适配真实设备的遗传锚点。设计跨模态联合编码器,通过对比学习对齐锚点与其对应视觉特征,生成鲁棒的联合表征,实现全片基因表达的高精度预测。实验表明,本方法在仅需10%转录组覆盖的情况下仍显著优于现有方法,且在无基因锚点时也表现更优。代码已开源。

原文摘要 · Abstract (English)

The high cost of spatial transcriptomics (ST) has driven extensive studies into predicting gene expression directly from H&E histology images. However, this prediction task faces an inherent limitation, as tissue morphology alone provides insufficient information to fully resolve underlying gene expression. To address this limitation, a recent study leverages partial gene expression to guide the prediction process alongside histology images. Building on this paradigm, we approach the prediction task as a spatial imputation problem, employing a Masked Autoencoder (MAE) to utilize a small fraction of gene expression as genetic anchors for inferring whole-slide gene expression profiles. Specifically, we propose a bio-saliency score and a learning-to-rank strategy to adaptively identify the most informative spots within the tissue. Based on these identified spots, our framework selects contiguous regions as genetic anchors to ensure suitability for real-world ST profiling hardware. To effectively leverage these anchors, we design a cross-modal joint encoder that integrates visual and genetic modalities. By aligning the selected anchors with their corresponding visual features via contrastive learning, the encoder generates robust joint representations to accurately predict gene expression across the whole slide. Notably, our framework consistently surpasses existing methods in both histology-only prediction and spatial imputation, achieving superior accuracy even without genetic anchors and further excelling with as little as 10% transcriptomic coverage. Our code is available at https://github.com/Kyyle2114/CAMMST.

基因预测多模态自编码器空间转录组

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。