arXiv:2605.07746stat.MLcs.LG2026-05

为计数数据设计新生成模型,提升生物数据建模效率与质量

Flow Matching for Count Data

论文配图:Flow Matching for Count Data
图 1 · 摘自论文原文
  • 基于连续时间出生-死亡过程构建计数数据生成框架
  • 无需模拟即可训练,参数量少且生成样本质量更高
  • 适用于单细胞测序和神经脉冲数据,路径可解释

高维计数数据广泛存在于单细胞RNA测序和神经脉冲序列等场景,跨批次或时间点的分布映射是数据分析的关键。受图像、视频和文本中扩散与流模型成功的启发,本文提出count-FM,一种基于局部单位跳变的连续时间出生-死亡过程的计数数据流匹配框架。该方法通过无模拟训练条件转移率,在计数空间中高效学习边际转移,实现任意计数分布源与目标之间的传输。实验表明,count-FM在生成样本质量上优于代表性基线,且参数量显著更少。进一步应用于scRNA-seq和神经脉冲数据,实现无条件生成、分布传输和条件生成,均取得更优样本质量、更高建模效率和可解释的传输路径。

原文摘要 · Abstract (English)

High-dimensional count data arise in applications such as single-cell RNA sequencing and neural spike trains, where mapping between distributions across successive batches or time points form critical components of data analysis. The recent success of diffusion- and flow-based deep generative models for images, video, and text motivates extending these ideas to count-valued settings, but many existing methods either treat each count as a categorical state or transform counts into a continuous space, neither of which is natural or efficient when the count range is large. We propose count-FM, a flow-matching framework for count data based on a continuous-time birth-death process with local unit jumps. Count-FM learns marginal transitions efficiently in count space through simulation-free training of conditional transition rates, allowing transport between arbitrary count-distributed source and target populations. In simulation, count-FM achieves better sample quality than representative baselines while using substantially fewer parameters. We further apply count-FM to scRNA-seq and neural spike-train data for unconditional generation, transport, and conditional generation. Across these tasks, count-FM yields improved sample quality, greater modeling efficiency, and interpretable transport paths.

生成模型计数数据生物信息流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。