arXiv:2603.07895cs.CV2026-03被引 1

用空间转录组数据增强病理模型,让其同时理解组织形态与分子状态。

MINT: Molecularly Informed Training with Spatial Transcriptomics Supervision for Pathology Foundation Models

  • 引入可学习的转录组标记,分离编码分子信息与形态信息。
  • 在577个HEST样本上训练,基因表达预测准确率达0.440(均Pearson r)。
  • 适合需要融合多模态病理数据的研究者,尤其关注分子机制的场景。

病理基础模型通过大规模全切片图像自监督预训练学习形态表征,但未显式捕捉组织的分子状态。空间转录组技术通过原位测量基因表达,提供了天然的跨模态监督信号。我们提出MINT(分子感知训练),一种将空间转录组监督融入预训练病理视觉变换器的微调框架。MINT在ViT输入中添加可学习的ST标记,独立编码转录组信息,避免灾难性遗忘;通过DINO自蒸馏和对冻结预训练编码器的显式特征锚定实现稳定优化。在斑点级(Visium)和补丁级(Xenium)分辨率下进行基因表达回归,提供跨空间尺度的互补监督。在577个公开HEST样本上训练后,MINT在HEST-Bench基因表达预测任务上达到均Pearson r = 0.440的最佳表现,在EVA通用病理任务上达到0.803,证明空间转录组监督能有效补充以形态为中心的自监督预训练。

原文摘要 · Abstract (English)

Pathology foundation models learn morphological representations through self-supervised pretraining on large-scale whole-slide images, yet they do not explicitly capture the underlying molecular state of the tissue. Spatial transcriptomics technologies bridge this gap by measuring gene expression in situ, offering a natural cross-modal supervisory signal. We propose MINT (Molecularly Informed Training), a fine-tuning framework that incorporates spatial transcriptomics supervision into pretrained pathology Vision Transformers. MINT appends a learnable ST token to the ViT input to encode transcriptomic information separately from the morphological CLS token, preventing catastrophic forgetting through DINO self-distillation and explicit feature anchoring to the frozen pretrained encoder. Gene expression regression at both spot-level (Visium) and patch-level (Xenium) resolutions provides complementary supervision across spatial scales. Trained on 577 publicly available HEST samples, MINT achieves the best overall performance on both HEST-Bench for gene expression prediction (mean Pearson r = 0.440) and EVA for general pathology tasks (0.803), demonstrating that spatial transcriptomics supervision complements morphology-centric self-supervised pretraining.

病理模型空间转录组多模态学习视觉变换器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。