提出计数桥模型,实现生物计数数据的精准建模与解卷积。
Count Bridges enable Modeling and Deconvolving Transcriptomic Data
- 基于整数上的随机桥过程,构建可精确计算的计数生成模型。
- 在多种指标上超越流匹配等基线,在计数分布拟合任务中表现领先。
- 适用于单细胞基因表达建模与批量测序数据解卷积,适合生物信息学研究者。
许多现代生物学检测技术(如RNA测序)产生的是反映检测分子数量的整数计数数据,但这些数据常因多细胞混合而缺乏单细胞分辨率。尽管扩散和流匹配等生成模型已被拓展至非欧几里得和离散场景,如何有效建模整数数据并系统解卷积聚合观测仍不明确。本文提出Count Bridges,一种定义在整数上的随机桥过程,提供扩散模型在计数数据上的精确、可计算类比,具备闭式条件分布,支持高效训练与采样。通过期望最大化风格的方法,直接从聚合测量中训练模型,并将单元级计数视为潜在变量。在整数分布匹配基准测试中,其性能优于流匹配与离散流匹配基线。进一步应用于两大生物问题:以核苷酸分辨率建模单细胞基因表达数据,用于批量RNA-seq解卷积;以及将多细胞空间转录组点分解为单细胞计数谱。该方法为跨尺度、多模态生物计数数据的生成建模与解卷积提供了原理性基础。
原文摘要 · Abstract (English)
Many modern biological assays, including RNA sequencing, yield integer-valued counts that reflect the number of molecules detected. These measurements are often not at the desired resolution: while the unit of interest is typically a single cell, many measurement technologies produce counts aggregated over sets of cells. Although recent generative frameworks such as diffusion and flow matching have been extended to non-Euclidean and discrete settings, it remains unclear how best to model integer-valued data or how to systematically deconvolve aggregated observations. We introduce Count Bridges, a stochastic bridge process on the integers that provides an exact, tractable analogue of diffusion-style models for count data, with closed-form conditionals for efficient training and sampling. We extend this framework to enable direct training from aggregated measurements via an Expectation-Maximization-style approach that treats unit-level counts as latent variables. We demonstrate state-of-the-art performance on integer distribution matching benchmarks, comparing against flow matching and discrete flow matching baselines across various metrics. We then apply Count Bridges to two large-scale problems in biology: modeling single-cell gene expression data at the nucleotide resolution, with applications to deconvolving bulk RNA-seq, and resolving multicellular spatial transcriptomic spots into single-cell count profiles. Our methods offer a principled foundation for generative modeling and deconvolution of biological count data across scales and modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。