构建首个标准化空间转录组数据基准,提升基因表达预测精度。
Completing Spatial Transcriptomics Data for Gene Expression Prediction Benchmarking
- 用Transformer模型补全低质量基因表达数据
- 新方法使误差降低超82.5%,显著优于现有技术
- 适合生物医学与深度学习交叉研究者参考
空间转录组学将组织切片图像与空间定位的基因表达数据结合,其中Visium技术应用最广,但受限于高成本、专业门槛和采样效率低导致的数据缺失。为此,研究者尝试从组织图像直接预测基因表达,但因数据集、预处理和训练方式不一致,模型比较困难。本文提出SpaRED数据库,整合26个公开数据集,建立统一评估标准;提出SpaCKLE模型,基于Transformer实现基因表达补全,相比现有方法平均误差降低82.5%以上;并通过该基准评估8个先进模型,验证其在原始与补全数据上的性能提升。本工作是目前最全面的空间转录组基因表达预测基准,为后续研究提供基础。
原文摘要 · Abstract (English)
Spatial Transcriptomics is a groundbreaking technology that integrates histology images with spatially resolved gene expression profiles. Among the various Spatial Transcriptomics techniques available, Visium has emerged as the most widely adopted. However, its accessibility is limited by high costs, the need for specialized expertise, and slow clinical integration. Additionally, gene capture inefficiencies lead to significant dropout, corrupting acquired data. To address these challenges, the deep learning community has explored the gene expression prediction task directly from histology images. Yet, inconsistencies in datasets, preprocessing, and training protocols hinder fair comparisons between models. To bridge this gap, we introduce SpaRED, a systematically curated database comprising 26 public datasets, providing a standardized resource for model evaluation. We further propose SpaCKLE, a state-of-the-art transformer-based gene expression completion model that reduces mean squared error by over 82.5% compared to existing approaches. Finally, we establish the SpaRED benchmark, evaluating eight state-of-the-art prediction models on both raw and SpaCKLE-completed data, demonstrating SpaCKLE substantially improves the results across all the gene expression prediction models. Altogether, our contributions constitute the most comprehensive benchmark of gene expression prediction from histology images to date and a stepping stone for future research on Spatial Transcriptomics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。