将自编码器升级为分布级表示学习,可捕捉多尺度生物数据中的统计特征。
Generative Distribution Embeddings: Lifting autoencoders to the space of distributions for multiscale representation learning
- 用生成器替代解码器,编码器处理样本集以学习分布嵌入。
- 在高斯与混合高斯分布上,潜空间距离近似恢复W2距离和最优传输路径。
- 适用于单细胞测序、基因调控、病毒蛋白演化等6个计算生物学难题。
许多真实世界问题需要跨尺度推理,要求模型作用于整个分布而非单个数据点。我们提出生成分布嵌入(GDE),将自编码器提升至分布空间。在GDE中,编码器作用于样本集合,解码器被生成器取代,目标是匹配输入分布。该框架通过条件生成模型与满足分布不变性准则的编码网络耦合,学习分布的表示。我们证明GDE能学习嵌入在Wasserstein空间中的预测充分统计量,潜空间距离近似恢复$W_2$距离,潜空间插值近似恢复高斯及高斯混合分布的最优传输轨迹。我们在合成数据集上系统评估,性能持续优于现有方法。进一步应用于六个计算生物学任务:从单核RNA测序数据(600万细胞)学习供体级表示,捕捉谱系追踪数据(15万细胞)中的克隆动态,预测扰动对转录组的影响(100万细胞),预测对细胞表型的影响(2000万单细胞图像),设计合成酵母启动子(3400万序列),以及病毒蛋白序列的时空建模(100万序列)。
原文摘要 · Abstract (English)
Many real-world problems require reasoning across multiple scales, demanding models which operate not on single data points, but on entire distributions. We introduce generative distribution embeddings (GDE), a framework that lifts autoencoders to the space of distributions. In GDEs, an encoder acts on sets of samples, and the decoder is replaced by a generator which aims to match the input distribution. This framework enables learning representations of distributions by coupling conditional generative models with encoder networks which satisfy a criterion we call distributional invariance. We show that GDEs learn predictive sufficient statistics embedded in the Wasserstein space, such that latent GDE distances approximately recover the $W_2$ distance, and latent interpolation approximately recovers optimal transport trajectories for Gaussian and Gaussian mixture distributions. We systematically benchmark GDEs against existing approaches on synthetic datasets, demonstrating consistently stronger performance. We then apply GDEs to six key problems in computational biology: learning donor-level representations from single-nuclei RNA sequencing data (6M cells), capturing clonal dynamics in lineage-traced RNA sequencing data (150K cells), predicting perturbation effects on transcriptomes (1M cells), predicting perturbation effects on cellular phenotypes (20M single-cell images), designing synthetic yeast promoters (34M sequences), and spatiotemporal modeling of viral protein sequences (1M sequences).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。