用生物信息学增强的攻击,发现基因生成模型存在隐私漏洞
Biologically-Informed Hybrid Membership Inference Attacks on Generative Genomic Models
- 结合语言模型与基因组上下文设计混合攻击方法
- 对小型和大型Transformer模型均实现更高成功率的隐私泄露
- 适合关注基因数据隐私保护的研究者和从业者
遗传数据的广泛可用性推动了基因组研究的发展,但也引发了敏感数据处理中的隐私担忧。本文探索使用语言模型(LM)生成合成基因突变谱型,并采用差分隐私(DP)保护敏感遗传数据。我们通过引入一种新型生物信息学启发的混合成员推断攻击(biHMIA),实证评估了DP模型的隐私保障效果。该攻击融合传统黑盒成员推断与基因组上下文指标,增强了攻击能力。实验表明,无论是小型还是大型Transformer类模型,均可有效生成小规模基因组数据;且本方法相较传统基于度量的成员推断攻击,在平均意义上实现了更高的攻击成功率。
原文摘要 · Abstract (English)
The increased availability of genetic data has transformed genomics research, but raised many privacy concerns regarding its handling due to its sensitive nature. This work explores the use of language models (LMs) for the generation of synthetic genetic mutation profiles, leveraging differential privacy (DP) for the protection of sensitive genetic data. We empirically evaluate the privacy guarantees of our DP modes by introducing a novel Biologically-Informed Hybrid Membership Inference Attack (biHMIA), which combines traditional black box MIA with contextual genomics metrics for enhanced attack power. Our experiments show that both small and large transformer GPT-like models are viable synthetic variant generators for small-scale genomics, and that our hybrid attack leads, on average, to higher adversarial success compared to traditional metric-based MIAs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。