用时变隐马尔可夫模型生成隐私保护的基因组数据,兼顾安全与真实度。
DP-SNP-TIHMM: Differentially Private, Time-Inhomogeneous Hidden Markov Models for Synthesizing Genome-Wide Association Datasets
- 基于时变隐马尔可夫模型生成全序列SNP数据,直接保护关联性隐私。
- 在ε∈[1,10]、δ=10⁻⁴下,合成数据保留原始统计特性。
- 适合需共享基因组数据的研究者,尤其关注隐私与实用性平衡者。
单核苷酸多态性(SNP)数据是遗传研究的基础,但共享时存在重大隐私风险。由于SNP间存在强相关性,可能导致掩码值重构、亲缘关系和成员身份推断等强力攻击。现有隐私保护方法要么对统计数据应用差分隐私,要么需要复杂后处理及公开数据集来抑制或选择性共享SNP。本研究提出一种创新框架,利用时变隐马尔可夫模型(TIHMM)生成合成SNP序列数据。通过确保每个SNP序列在训练中仅贡献有限影响,实现强差分隐私保障。关键在于,通过对完整序列及其梯度贡献进行边界控制,直接应对由序列相关性引发的隐私风险。在真实世界1000 Genomes数据集上实验表明,在ε∈[1,10]、δ=10⁻⁴的隐私预算下,该方法有效生成高保真合成数据。通过允许马尔可夫转移模型依赖于序列位置,显著提升性能,使合成数据能高度复现非私有数据的统计特征。该框架实现了基因组数据的安全共享,同时为研究者提供极高的灵活性与可用性。
原文摘要 · Abstract (English)
Single nucleotide polymorphism (SNP) datasets are fundamental to genetic studies but pose significant privacy risks when shared. The correlation of SNPs with each other makes strong adversarial attacks such as masked-value reconstruction, kin, and membership inference attacks possible. Existing privacy-preserving approaches either apply differential privacy to statistical summaries of these datasets or offer complex methods that require post-processing and the usage of a publicly available dataset to suppress or selectively share SNPs. In this study, we introduce an innovative framework for generating synthetic SNP sequence datasets using samples derived from time-inhomogeneous hidden Markov models (TIHMMs). To preserve the privacy of the training data, we ensure that each SNP sequence contributes only a bounded influence during training, enabling strong differential privacy guarantees. Crucially, by operating on full SNP sequences and bounding their gradient contributions, our method directly addresses the privacy risks introduced by their inherent correlations. Through experiments conducted on the real-world 1000 Genomes dataset, we demonstrate the efficacy of our method using privacy budgets of $\varepsilon \in [1, 10]$ at $δ=10^{-4}$. Notably, by allowing the transition models of the HMM to be dependent on the location in the sequence, we significantly enhance performance, enabling the synthetic datasets to closely replicate the statistical properties of non-private datasets. This framework facilitates the private sharing of genomic data while offering researchers exceptional flexibility and utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。