arXiv:2602.03477cs.LGcs.AI2026-02被引 1

用掩码扩散模型同时生成单细胞基因身份和表达值,避免顺序偏差。

ScDiVa: Masked Discrete Diffusion for Joint Modeling of Single-Cell Identity and Expression

  • 通过连续时间掩码机制模拟数据缺失,实现双向去噪联合建模。
  • 在5900万细胞上预训练,跨批次整合与扰动预测性能优异。
  • 适合需要高质量单细胞生成的生物医学研究者使用。

单细胞RNA测序数据维度高、稀疏且无序,传统自回归生成方法会引入人为排序偏差并积累误差。为此,我们提出scDiVa,一种基于掩码离散扩散的奠基模型,通过在词元空间定义连续时间前向掩码机制,使生成过程与类似丢失的噪声过程对齐。scDiVa采用双向去噪器,联合建模离散基因身份与连续表达值,利用熵归一化序列化与潜在锚点词元,提升信息效率并保持全局细胞身份。模型通过深度无关的时间采样与双重去噪目标进行训练,以模拟不同稀疏水平并精确恢复身份与数值。在5900万细胞上预训练后,scDiVa在主流基准测试中表现强劲,涵盖批次整合、细胞类型注释及扰动响应预测。结果表明,掩码离散扩散为自回归提供了一种生物学合理且高效的替代方案。

原文摘要 · Abstract (English)

Single-cell RNA-seq profiles are high-dimensional, sparse, and unordered, causing autoregressive generation to impose an artificial ordering bias and suffer from error accumulation. To address this, we propose scDiVa, a masked discrete diffusion foundation model that aligns generation with the dropout-like corruption process by defining a continuous-time forward masking mechanism in token space. ScDiVa features a bidirectional denoiser that jointly models discrete gene identities and continuous values, utilizing entropy-normalized serialization and a latent anchor token to maximize information efficiency and preserve global cell identity. The model is trained via depth-invariant time sampling and a dual denoising objective to simulate varying sparsity levels while ensuring precise recovery of both identity and magnitude. Pre-trained on 59 million cells, scDiVa achieves strong transfer performance across major benchmarks, including batch integration, cell type annotation, and perturbation response prediction. These results suggest that masked discrete diffusion serves as a biologically coherent and effective alternative to autoregression.

单细胞生成扩散模型生物信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。