用大模型知识压缩单细胞数据,生成更小更准的合成数据集。
scDD: Latent Codes Based scRNA-seq Dataset Distillation with Foundation Model Knowledge
- 将原始单细胞数据与大模型知识融合到隐空间,生成合成数据。
- 单步扩散生成器提升优化效率,保持细胞类型区分度和数据特征。
- 在多任务评估中平均性能超越现有方法15.7%以上,适合数据共享场景。
单细胞RNA测序(scRNA-seq)已对数亿个人类细胞进行了跨器官、疾病、发育及扰动的分析。然而,原始数据存在高维稀疏性、批次效应噪声、类别不平衡以及持续增长的数据规模,给多中心知识迁移、数据融合和跨数据集验证带来挑战。为此,我们提出基于隐码的scRNA-seq数据蒸馏框架scDD,将基础模型知识与原始数据信息映射至紧凑隐空间,并通过生成器生成替代原始数据的合成scRNA-seq数据集。进一步提出单步条件扩散生成器SCDG,通过单步梯度反向传播优化蒸馏质量,避免多步传播导致的梯度衰减;同时借助灵活的条件控制与生成质量保障机制,确保合成数据保留真实数据特征及类间可区分性。最后构建全面基准,评估不同数据分析任务中的蒸馏性能。结果表明,本方法在平均任务上相较先前最优方法实现7.61%绝对提升和15.70%相对提升。
原文摘要 · Abstract (English)
Single-cell RNA sequencing (scRNA-seq) technology has profiled hundreds of millions of human cells across organs, diseases, development and perturbations to date. However, the high-dimensional sparsity, batch effect noise, category imbalance, and ever-increasing data scale of the original sequencing data pose significant challenges for multi-center knowledge transfer, data fusion, and cross-validation between scRNA-seq datasets. To address these barriers, (1) we first propose a latent codes-based scRNA-seq dataset distillation framework named scDD, which transfers and distills foundation model knowledge and original dataset information into a compact latent space and generates synthetic scRNA-seq dataset by a generator to replace the original dataset. Then, (2) we propose a single-step conditional diffusion generator named SCDG, which perform single-step gradient back-propagation to help scDD optimize distillation quality and avoid gradient decay caused by multi-step back-propagation. Meanwhile, SCDG ensures the scRNA-seq data characteristics and inter-class discriminability of the synthetic dataset through flexible conditional control and generation quality assurance. Finally, we propose a comprehensive benchmark to evaluate the performance of scRNA-seq dataset distillation in different data analysis tasks. It is validated that our proposed method can achieve 7.61% absolute and 15.70% relative improvement over previous state-of-the-art methods on average task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。