arXiv:2603.10873cs.LGq-bio.GN2026-03被引 1

用表型监督生成真实基因型数据,保护隐私同时保持预测能力。

SNPgen: Phenotype-Supervised Genotype Representation and Synthetic Data Generation via Latent Diffusion

  • 两阶段扩散模型,结合GWAS选关键位点和条件生成
  • 合成数据在4种疾病上达到真实数据的预测性能
  • 适合需要隐私保护的基因组研究团队使用

多基因风险评分等基因组分析需要大规模个体基因型数据,但严格的数据访问限制阻碍了数据共享。合成基因型生成可提供隐私保护替代方案,但现有方法多为无条件生成,缺乏表型对齐,或依赖无监督压缩,导致统计保真度与下游任务效用之间存在差距。我们提出SNPgen,一种两阶段的条件潜在扩散框架,用于生成表型监督的合成基因型。SNPgen结合GWAS引导的变异选择(1,024-2,048个与性状相关SNPs),通过变分自编码器进行基因型压缩,并利用分类器自由引导的潜在扩散模型,以二元疾病标签为条件生成。在包含458,724名英国生物银行个体的四种复杂疾病(冠心病、乳腺癌、1型和2型糖尿病)上评估,基于合成数据训练的模型在“用合成数据训练、真实数据测试”协议下,预测性能接近使用2-6倍更多位点的全基因组PRs方法。隐私分析显示零完全匹配,近随机的成员推断(AUC≈0.50),保留了连锁不平衡结构,且等位基因频率相关性高(r≥0.95)。通过已知因果效应的受控模拟验证了所施加遗传关联结构的忠实恢复。

原文摘要 · Abstract (English)

Polygenic risk scores and other genomic analyses require large individual-level genotype datasets, yet strict data access restrictions impede sharing. Synthetic genotype generation offers a privacy-preserving alternative, but most existing methods operate unconditionally, producing samples without phenotype alignment, or rely on unsupervised compression, creating a gap between statistical fidelity and downstream task utility. We present SNPgen, a two-stage conditional latent diffusion framework for generating phenotype-supervised synthetic genotypes. SNPgen combines GWAS-guided variant selection (1,024-2,048 trait-associated SNPs) with a variational autoencoder for genotype compression and a latent diffusion model conditioned on binary disease labels via classifier-free guidance. Evaluated on 458,724 UK Biobank individuals across four complex diseases (coronary artery disease, breast cancer, type 1 and type 2 diabetes), models trained on synthetic data matched real-data predictive performance in a train-on-synthetic, test-on-real protocol, approaching genome-wide PRS methods that use $2$-$6\times$ more variants. Privacy analysis confirmed zero identical matches, near-random membership inference (AUC $\approx 0.50$), preserved linkage disequilibrium structure, and high allele frequency correlation ($r \geq 0.95$) with source data. A controlled simulation with known causal effects verified faithful recovery of the imposed genetic association structure.

基因组生成扩散模型隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。