arXiv:2609.05956cs.CV2026-09

构建首个统一虚拟空间转录组基准,提升模型评估可靠性。

STP-BENCH: A Unified Systematic Benchmark for Virtual Spatial Transcriptomics from Histopathology Images

论文配图:STP-BENCH: A Unified Systematic Benchmark for Virtual Spatial Transcriptomics from Histopathology Images
图 1 · 摘自论文原文
  • 设计标准化基准,整合六种癌症、两大平台数据
  • 21个模型统一评估,发现编码方式显著影响排名
  • 强调生物可解释性与跨域鲁棒性,适合算法与病理研究者

空间转录组(ST)能揭示肿瘤异质性,但实验成本高。因此,从苏木精-伊红染色切片预测空间基因表达的虚拟ST方法迅速发展。然而,现有评估存在数据小、不一致、缺乏生物学可解释性等问题。为此,我们提出STP-BENCH,涵盖六种癌症类型、两个ST平台(Visium和Xenium),每个训练集含超过30,000个点且至少15张切片,确保统计可靠性。我们评估了21种预测方法,并在架构允许时统一使用病理基础模型作为形态编码器。不仅报告平均预测精度,还系统分析哪些基因和基因集可从组织形态中恢复,进一步通过细胞类型去卷积和空间区域识别评估预测谱的下游生物学价值,并测试模型在领域迁移和数据缩放下的可靠性。结果表明,统一编码显著改变以往模型排名,说明此前评估中架构创新与图像编码被混淆。我们公开发布STP-BENCH,支持可复现研究,网址:https://github.com/NEXGEM/STP-Bench。

原文摘要 · Abstract (English)

Spatial transcriptomics (ST) provides unprecedented insights into tumor heterogeneity by capturing spatially resolved gene expression, yet its high experimental cost hinders large-scale adoption. Consequently, computational approaches that predict spatial gene expression directly from hematoxylin and eosin slides, termed virtual ST, have rapidly emerged. Despite this progress, assessing advances in the field remains difficult due to insufficient benchmarking: prior studies rely on small, heterogeneous datasets, inconsistent training and inference pipelines, and limited evaluation of biological interpretability and model robustness. To address these gaps, we present STP-BENCH, a standardized benchmark for virtual ST models. STP-BENCH comprises six cancer types spanning two ST platforms (Visium and Xenium), with each training dataset containing more than 30,000 spots and at least 15 slides to ensure statistical reliability. We evaluate 21 predictive approaches, re-implemented with a unified pathology foundation model as the morphological encoder when architecturally applicable. Beyond conventional benchmarks that report average predictive accuracy on highly variable genes, we systematically examine which genes and gene sets are recoverable from histomorphology. We further evaluate the downstream biological utility of predicted profiles through cell-type deconvolution and spatial domain identification, and assess model reliability under domain shifts and data scaling. Notably, unified morphological encoding substantially re-orders model rankings established in prior studies, indicating that architectural innovations and image encoding have been conflated in previous evaluations. We publicly release STP-BENCH to support reproducibility and serve as a community benchmark at https://github.com/NEXGEM/STP-Bench.

空间转录组虚拟检测医学图像基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。