arXiv:2505.07188cs.CRcs.CL2025-05被引 5

针对联邦学习中基因组数据的推理攻击,提出防护机制研究

Securing Genomic Data Against Inference Attacks in Federated Learning Environments

  • 用合成基因组数据模拟联邦学习环境,测试三类推理攻击
  • 梯度型成员推断攻击精度达0.79,F1-score为0.87,风险最高
  • 揭示原始联邦学习无法保护基因隐私,适合生物信息与隐私安全研究者

联邦学习(FL)为在去中心化基因组数据集上协作训练机器学习模型提供了前景广阔的框架,无需直接共享数据。尽管该方法保持了数据本地性,但仍易受复杂推理攻击,威胁个体隐私。本研究使用合成基因组数据模拟联邦学习场景,评估三类关键攻击向量:成员推断攻击(MIA)、基于梯度的成员推断攻击和标签推断攻击(LIA)。实验显示,基于梯度的MIA效果最佳,精度为0.79,F1-score达0.87,凸显梯度暴露在联邦更新中的风险。通过雷达图对比攻击性能,并量化各客户端的模型泄露程度。结果表明,简单联邦学习设置不足以保障基因组隐私,亟需针对基因数据高度敏感特性设计更鲁棒的隐私保护机制。

原文摘要 · Abstract (English)

Federated Learning (FL) offers a promising framework for collaboratively training machine learning models across decentralized genomic datasets without direct data sharing. While this approach preserves data locality, it remains susceptible to sophisticated inference attacks that can compromise individual privacy. In this study, we simulate a federated learning setup using synthetic genomic data and assess its vulnerability to three key attack vectors: Membership Inference Attack (MIA), Gradient-Based Membership Inference Attack, and Label Inference Attack (LIA). Our experiments reveal that Gradient-Based MIA achieves the highest effectiveness, with a precision of 0.79 and F1-score of 0.87, underscoring the risk posed by gradient exposure in federated updates. Additionally, we visualize comparative attack performance through radar plots and quantify model leakage across clients. The findings emphasize the inadequacy of naïve FL setups in safeguarding genomic privacy and motivate the development of more robust privacy-preserving mechanisms tailored to the unique sensitivity of genomic data.

联邦学习基因组隐私推理攻击隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。