基于表格结构与联邦学习的单细胞基因调控模型,实现隐私保护下的衰老研究。
Predictive single cell foundation model for gene regulation and aging with privacy-preserving tabular learning
- 显式建模单细胞数据的表格结构,结合联邦学习保护隐私。
- 在多类生物系统中发现组合调控规律,识别出潜在抗衰老因子。
- 适合生物医学研究者、隐私敏感领域团队使用。
预训练基础模型正在改变单细胞基因组学,但其扩展引发隐私担忧。与文本数据不同,单细胞数据无序且具有独特表格结构,现有模型未予考虑。我们提出Tabula,一种基于联邦学习(FL)的隐私保护基础模型,显式建模单细胞数据的表格结构。为部署Tabula,我们进一步开发了Chiron平台,支持机构间去中心化协作训练,无需共享原始数据。除在下游任务中表现优异外,Tabula揭示了造血、胰腺发育、神经发生和心肌生成等多样生物系统中的组合调控逻辑。利用新的人类成纤维细胞年轻与老年配对scRNA-seq数据集,通过年龄与身份得分引导的体外筛选,成功提名抗衰老因子,优于传统方法。因此,Tabula通过融合表格学习与联邦学习,推动了隐私保护的单细胞基础建模发展,为人类健康提供隐私安全的虚拟细胞范式。
原文摘要 · Abstract (English)
Pre-trained foundation models (FMs) have begun transforming single-cell genomics, but scaling them raises privacy concerns. Moreover, unlike text data, single-cell data is unordered and exhibits a unique tabular structure that current single-cell FMs overlook. We introduce Tabula, a privacy-preserving FM designed with federated learning (FL) that explicitly models the tabular structure of single-cell data. To deploy Tabula, we further developed Chiron, a decentralized AI agent-enabled platform for collaborative training across institutions without sharing raw data. Beyond strong performance across downstream benchmarks, Tabula reveals combinatorial regulatory logic across diverse biological systems, including hematopoiesis, pancreatic endogenesis, neurogenesis, and cardiogenesis. Using a new scRNA-seq dataset of paired young and aged human fibroblasts, Tabula nominates rejuvenation factors through age- and identity score-guided in silico prioritization, outperforming conventional approaches. Thus, Tabula represents an important advance in single-cell foundation modeling by integrating tabular learning with FL, paving the way toward privacy-preserving virtual cells for human health.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。