arXiv:2603.28334cs.LGcs.DC2026-03

用嵌入密钥的轻量联邦学习保护生物组学数据隐私

Key-Embedded Privacy for Decentralized AI in Biomedical Omics

  • 将秘密密钥直接嵌入神经网络结构,实现可控制的隐私保护
  • 在多种组学任务中保持高性能,误差率低于基准方法15%以上
  • 适合需要跨机构协作但数据敏感的生物医药研究团队

生物医学领域数据驱动方法的快速发展加剧了隐私、治理与监管方面的担忧,限制了原始数据共享,阻碍了具有临床意义的代表性队列构建。现有加密方案通常开销过大,差分隐私又会降低模型性能,导致实际应用效果不佳。为此,我们提出一种基于隐式神经表示的轻量级联邦学习方法INFL,通过在客户端模型中集成即插即用的坐标条件模块,将秘密密钥直接嵌入网络架构,并支持异构站点间的无缝聚合。在多种生物组学任务中——包括大规模蛋白质组学分类、单细胞转录组学扰动预测回归,以及空间转录组学和多组学的聚类分析(含公开与私有数据)——我们验证了INFL在保障强且可控隐私的同时,维持了足够的模型效用,性能足以支撑下游科研与临床应用。

原文摘要 · Abstract (English)

The rapid adoption of data-driven methods in biomedicine has intensified concerns over privacy, governance, and regulation, limiting raw data sharing and hindering the assembly of representative cohorts for clinically relevant AI. This landscape necessitates practical, efficient privacy solutions, as cryptographic defenses often impose heavy overhead and differential privacy can degrade performance, leading to sub-optimal outcomes in real-world settings. Here, we present a lightweight federated learning method, INFL, based on Implicit Neural Representations that addresses these challenges. Our approach integrates plug-and-play, coordinate-conditioned modules into client models, embeds a secret key directly into the architecture, and supports seamless aggregation across heterogeneous sites. Across diverse biomedical omics tasks, including cohort-scale classification in bulk proteomics, regression for perturbation prediction in single-cell transcriptomics, and clustering in spatial transcriptomics and multi-omics with both public and private data, we demonstrate that INFL achieves strong, controllable privacy while maintaining utility, preserving the performance necessary for downstream scientific and clinical applications.

联邦学习隐私保护生物组学密钥嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。