arXiv:2509.01864cs.CV2025-09ICCV

无需参考数据,用扩散模型填补空间转录组缺失基因表达

Latent Gene Diffusion for Spatial Transcriptomics Completion

  • 提出无参考的潜变量基因扩散模型,利用上下文基因构建生物有意义的基因潜空间
  • 在26个数据集上平均均方误差降低18%,且提升其他6种方法性能最高10%
  • 适合处理基因表达缺失问题的研究者,尤其关注生物可解释性的建模场景

计算机视觉在分析空间转录组(ST)数据方面已展现出强大能力。然而,当前基于组织病理图像预测空间分辨基因表达的模型受限于数据缺失问题。多数现有方法依赖单细胞RNA测序参考数据,受配准质量与外部数据集影响,易引入批次效应和继承性缺失。本文提出LGDiST,首个无需参考的潜变量基因扩散模型,用于解决ST数据缺失问题。实验表明,LGDiST在26个数据集上的平均均方误差比先前最优方法低18%;使用LGDiST完成数据后,六种先进方法的基因表达预测性能在均方误差上最高提升10%。其关键创新在于利用以往被认为信息量低的上下文基因,构建丰富且具有生物学意义的基因潜空间。实验显示,移除上下文基因、ST潜空间或邻居条件等组件均导致性能显著下降,证明完整架构优于各组件独立表现。

原文摘要 · Abstract (English)

Computer Vision has proven to be a powerful tool for analyzing Spatial Transcriptomics (ST) data. However, current models that predict spatially resolved gene expression from histopathology images suffer from significant limitations due to data dropout. Most existing approaches rely on single-cell RNA sequencing references, making them dependent on alignment quality and external datasets while also risking batch effects and inherited dropout. In this paper, we address these limitations by introducing LGDiST, the first reference-free latent gene diffusion model for ST data dropout. We show that LGDiST outperforms the previous state-of-the-art in gene expression completion, with an average Mean Squared Error that is 18% lower across 26 datasets. Furthermore, we demonstrate that completing ST data with LGDiST improves gene expression prediction performance on six state-of-the-art methods up to 10% in MSE. A key innovation of LGDiST is using context genes previously considered uninformative to build a rich and biologically meaningful genetic latent space. Our experiments show that removing key components of LGDiST, such as the context genes, the ST latent space, and the neighbor conditioning, leads to considerable drops in performance. These findings underscore that the full architecture of LGDiST achieves substantially better performance than any of its isolated components.

空间转录组基因表达扩散模型数据补全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。