arXiv:2605.16392q-bio.QMcs.CV2026-05被引 1

用基因数据生成虚拟染色,让病理图像分析更懂分子信息。

Bridging the Modality Bottleneck in Pathology MIL through Virtual Molecular Staining

论文配图:Bridging the Modality Bottleneck in Pathology MIL through Virtual Molecular Staining
图 1 · 摘自论文原文
  • 用配对的转录组数据训练虚拟分子染色,重构病理图像特征。
  • 在23个任务中平均提升3.5%,生存预测最高增5.2%。
  • 无需基因数据即可推理,适合各类病理分析模型使用。

多实例学习(MIL)是计算病理学中全切片图像分析的主流框架,通常由冻结的图像块编码器、投影层和整体切片聚合器组成。尽管编码器和聚合器已广泛研究,投影层仍存在仅依赖形态的瓶颈,难以捕捉与生物标志物状态和生存率相关的分子信息。本文提出分子感知染色变换(MIST),作为MIL投影层的即插即用替代方案。MIST仅在训练阶段使用配对的空间转录组数据,构建虚拟分子染色。它将基因表达谱聚类为跨模态原型,将其锚定在冻结的基础模型特征空间中,并据此重新组织H&E图像块特征。推理阶段无需转录组数据,可插入标准MIL聚合器前。我们在23个下游任务和8种MIL聚合器上评估MIST,结果表明其在256种配置中改进240种,平均提升3.5%;各项任务均有显著增益:生存预测+5.2%,组织亚型分类+3.3%,生物标志物预测+2.6%。消融实验确认基因原型是性能提升主因,空间、生物学及病理分析显示跨模态原型亲和力能从H&E图像中捕捉空间一致的分子程序。

原文摘要 · Abstract (English)

Multiple instance learning (MIL) is the dominant framework for whole-slide image analysis in computational pathology, typically combining a frozen patch encoder, a projection layer, and a slide-level aggregator. While encoders and aggregators have been extensively studied, the projection layer remains a largely morphology-only bottleneck. This limits endpoints such as biomarker status and survival, which are governed by a molecular state that is not fully captured by H&E morphology. We introduce Molecularly Informed Staining Transform (MIST), a plug-in replacement for the MIL projection layer that uses paired spatial transcriptomics only during training to construct virtual molecular stains. MIST clusters gene expression profiles into cross-modal prototypes, anchors them in the frozen foundation model feature space, and uses them to reorganize H&E patch features along molecularly guided axes. It requires no transcriptomics at inference and can be inserted before standard MIL aggregators. We evaluate MIST across 23 downstream tasks and 8 MIL aggregators. MIST improves 240 of 256 configurations over the standard projection layer, with an average gain of +3.5%, observed consistently across endpoint types: +5.2% on survival prediction, +3.3% on tissue subtyping, and +2.6% on biomarker prediction. Ablations confirm that gene-derived prototypes are the primary source of the gains, while spatial, biological, and pathological analyses show that cross-modal prototype affinities capture spatially coherent molecular programs from H&E alone.

病理图像多模态分子染色深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。