用扩散模型生成更真实的单细胞基因表达数据,不依赖基因顺序。
Scalable Single-Cell Gene Expression Generation with Latent Diffusion Models
- 基于可交换性的潜空间扩散模型,避免人为设定基因顺序。
- 在观测与扰动数据上生成质量更高,分类任务表现更优。
- 适合生物信息学研究者,尤其关注真实表达数据生成者。
单细胞基因表达的计算建模对理解细胞过程至关重要,但生成真实表达谱仍面临挑战,源于基因表达数据的计数特性及基因间复杂的潜在依赖关系。现有生成模型常引入人工基因排序或依赖浅层神经网络结构。本文提出一种可扩展的潜空间扩散模型(scLDM),尊重数据的固有可交换性。其变分自编码器采用固定大小的潜变量,并利用统一的多头交叉注意力块(MCAB)架构,在编码器中实现排列不变性池化,在解码器中实现排列等变性还原。进一步通过扩散变换器与线性插值替换高斯先验,实现高质量生成及多条件无分类器引导。实验表明,该模型在观测与扰动单细胞数据上均表现优异,并提升下游细胞分类任务性能。
原文摘要 · Abstract (English)
Computational modeling of single-cell gene expression is crucial for understanding cellular processes, but generating realistic expression profiles remains a major challenge. This difficulty arises from the count nature of gene expression data and complex latent dependencies among genes. Existing generative models often impose artificial gene orderings or rely on shallow neural network architectures. We introduce a scalable latent diffusion model for single-cell gene expression data, which we refer to as scLDM, that respects the fundamental exchangeability property of the data. Our VAE uses fixed-size latent variables leveraging a unified Multi-head Cross-Attention Block (MCAB) architecture, which serves dual roles: permutation-invariant pooling in the encoder and permutation-equivariant unpooling in the decoder. We enhance this framework by replacing the Gaussian prior with a latent diffusion model using Diffusion Transformers and linear interpolants, enabling high-quality generation with multi-conditional classifier-free guidance. We show its superior performance in a variety of experiments for both observational and perturbational single-cell data, as well as downstream tasks like cell-level classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。