arXiv:2608.08366cs.CVq-bio.GN2026-08

用病理图像预测单细胞基因表达,成本低且效果优于现有方法。

VOICE: A Vision-Omics Foundation Model Integrating Direct and Retrieval-Based Prediction of In-situ Single-Cell Gene Expression

论文配图:VOICE: A Vision-Omics Foundation Model Integrating Direct and Retrieval-Based Prediction of In-situ Single-Cell Gene Expression
图 1 · 摘自论文原文
  • 融合形态与转录数据,双分支预测基因表达
  • 在7项指标上超越已有方法,泛化能力强
  • 适合大规模组织样本的分子分析与临床研究

空间转录组学可实现单细胞分辨率的基因表达解析,但成本高,仅限数百至数千个靶向基因,适用样本量小。相比之下,H&E染色图像成本低、常规大量采集。因此,直接从形态预测单细胞表达是将分子分析扩展至大型组织存档的可行方案。我们提出 VOICE,一种多模态基础模型,利用配对的 Xenium 数据,从 H&E 图像预测单细胞基因表达。VOICE 首先通过对比学习在2300万细胞上训练,将病理基础模型的细胞中心形态与转录组基础模型的表达嵌入对齐。随后通过两个分支预测:一个直接从形态回归表达;另一个通过检索相似参考细胞,恢复无形态信号的基因。由于不同基因的形态可预测性差异,VOICE 使用每基因加权融合两分支。训练后,VOICE 在未见患者、切片及部分重叠基因面板上表现良好,七项指标均优于现有方法。

原文摘要 · Abstract (English)

Spatial transcriptomics can resolve gene expression at single-cell resolution, but it is costly, limited to targeted panels of a few hundred to a few thousand genes, and applicable to only a small number of samples. H&E imaging, by contrast, is cheap and collected routinely at scale. This makes predicting single-cell expression directly from morphology a practical way to bring molecular analysis to large tissue archives. We therefore present VOICE, a multimodal foundation model that predicts single-cell gene expression from H&E images using paired Xenium data. VOICE first aligns cell centered H&E morphology from a pathology foundation model with single-cell expression embeddings from a transcriptome foundation model, trained using contrastive learning over 23 million cells. Next it predicts expression through two branches. One branch directly regresses expression from morphology. The other branch retrieves measured expression from similar reference cells, recovering genes that do not have morphological signal. Because genes vary in morphological predictability, VOICE fuses the two branches with a per-gene weight. After training, VOICE generalizes to heldout patients, slides, and partially overlapping gene panels from Xenium, and it consistently outperforms prior single-cell expression prediction methods on seven metrics.

基因表达预测多模态模型空间转录组H&E图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。