针对单细胞表观基因组数据隐私与高维稀疏难题,提出首个适配的联邦学习框架。
FL-Sailer: Efficient and Privacy-Preserving Federated Learning for Scalable Single-Cell Epigenetic Data Analysis via Adaptive Sampling

- 通过自适应采样降低80%维度,保留生物意义特征
- 用不变变分自编码器分离生物信号与技术噪声
- 适合多机构协作研究,尤其擅长处理异质性数据
单细胞ATAC-seq(scATAC-seq)可实现染色质开放性的高分辨率图谱绘制,但隐私法规和数据规模限制了多机构间的数据共享。联邦学习(FL)提供了一种隐私保护的替代方案,但在scATAC-seq分析中面临三大挑战:超高维度、极端稀疏性以及严重的跨机构异质性。我们提出FL-Sailer,首个专为scATAC-seq数据设计的联邦学习框架。该框架融合两项核心创新:(i) 自适应杠杆得分采样,选择具有生物学意义的特征,将维度降低80%;(ii) 不变变分自编码器架构,通过最小化互信息来解耦生物信号与技术混杂因子。我们提供了收敛性保证,证明FL-Sailer可在有界误差内收敛至原高维问题的近似解。在合成与真实表观基因组数据集上的大量实验表明,FL-Sailer不仅实现了此前无法完成的多机构协作,还通过自适应采样作为隐式正则化抑制技术噪声,性能优于集中式方法。本研究证实,当联邦学习针对特定领域挑战进行定制时,可成为协作表观基因组研究的更优范式。
原文摘要 · Abstract (English)
Single-cell ATAC-seq (scATAC-seq) enables high-resolution mapping of chromatin accessibility, yet privacy regulations and data size constraints hinder multi-institutional sharing. Federated learning (FL) offers a privacy-preserving alternative, but faces three fundamental barriers in scATAC-seq analysis: ultra-high dimensionality, extreme sparsity, and severe cross-institutional heterogeneity. We propose FL-Sailer, the first FL framework designed for scATAC-seq data. FL-Sailer integrates two key innovations: (i) adaptive leverage score sampling, which selects biologically interpretable features while reducing dimensionality by 80%, and (ii) an invariant VAE architecture, which disentangles biological signals from technical confounders via mutual information minimization. We provide a convergence guarantee, showing that FL-Sailer converges to an approximate solution of the original high-dimensional problem with bounded error. Extensive experiments on synthetic and real epigenomic datasets demonstrate that FL-Sailer not only enables previously infeasible multi-institutional collaborations but also surpasses centralized methods by leveraging adaptive sampling as an implicit regularizer to suppress technical noise. Our work establishes that federated learning, when tailored to domain-specific challenges, can become a superior paradigm for collaborative epigenomic research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。