arXiv:2605.22340cs.LG2026-05

用流匹配建模单细胞基因表达动态,填补时间空白。

From Snapshots to Trajectories: Learning Single-Cell Gene Expression Dynamics via Conditional Flow Matching

论文配图:From Snapshots to Trajectories: Learning Single-Cell Gene Expression Dynamics via Conditional Flow Matching
图 1 · 摘自论文原文
  • 基于条件流匹配,利用熵正则化最优传输构建软目标
  • 双向速度场一致性和分布对齐提升长期预测稳定性
  • 适合研究细胞分化、发育等稀疏时间数据的轨迹重建

单细胞RNA测序(scRNA-seq)提供细胞状态的高维谱图,支持数据驱动的细胞动态建模。实际中,时间分辨的scRNA-seq仅在少数离散时间点采集为未配对的快照群体,导致显著的时间空隙。这促使对未测量时间点的轨迹推断。现有方法主要沿两个方向:最优传输(OT)对齐实现观测快照间的分布级匹配,而连续时间生成模型通过学习动态实现预测。然而仍面临两大挑战:(i) 未配对快照使相邻时间点间的局部转移模糊,导致监督不稳定;(ii) 长时程预测依赖重复积分,小的建模误差累积引发分布漂移。为此,我们提出单细胞流匹配(scFM),一种基于耦合条件流匹配的潜在生成框架。首先,计算相邻快照间的熵正则化OT耦合并用于构建软、加权的流匹配目标以学习时间依赖的速度场。其次,学习双向速度场并利用其一致性来优化耦合并提升稀疏监督下的时间一致性。第三,引入分布级对齐和潜在动态正则化以锚定长滚动并缓解漂移。在真实世界的时间序列scRNA-seq数据集上的实验表明,scFM在时间插值和外推任务中均显著提升分布预测性能。此外,scFM能更准确地重建轨迹并生成无中间时间点时也具时间一致性的可视化结果,表明其更忠实恢复了底层基因表达动态。

原文摘要 · Abstract (English)

Single-cell RNA sequencing (scRNA-seq) provides high-dimensional profiles of cellular states, enabling data-driven modeling of cellular dynamics over time. In practice, time-resolved scRNA-seq is collected at only a few discrete time points as unpaired snapshot populations, leaving substantial temporal gaps. This motivates trajectory inference at unmeasured time points. Existing methods mainly follow two directions, optimal-transport (OT) alignment provides distribution-level matching between observed snapshots, while continuous-time generative models support forecasting via learned dynamics. However, two challenges remain: (i) unpaired snapshots render local transitions between adjacent time points ambiguous, leading to unstable supervision; and (ii) long-horizon prediction relies on repeated integration, where small modeling errors compound and cause distribution drift. To address these challenges, we propose single-cell Flow Matching (scFM), a latent generative framework based on coupling-conditioned flow matching. First, we compute entropically regularized OT couplings between adjacent snapshots and use them to construct soft, weighted flow-matching targets for learning time-dependent velocity fields. Second, we learn bidirectional velocity fields and leverage their consistency to refine couplings and improve temporal coherence under sparse supervision. Third, we introduce distribution-level alignment and latent dynamic regularization to anchor long rollouts and mitigate drift. Experiments on real-world time-series scRNA-seq datasets show that scFM consistently improves distributional prediction performance for both temporal interpolation and extrapolation. Moreover, scFM yields more accurate trajectory reconstruction and temporally coherent visualizations where intermediate time points are absent, indicating a more faithful recovery of underlying temporal gene expression dynamics.

单细胞轨迹推断流匹配动态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。