arXiv:2507.05247cs.LG2025-07

提出无数据泄露的多病种深度学习框架,提升基因组关联分析精度。

Multi-Disease Deep Learning Framework for GWAS: Beyond Feature Selection Constraints

  • 设计无特征筛选泄漏的深度网络架构,避免路径依赖
  • 在37,000样本、500万SNP上实现AUC 0.68-0.96
  • 适用于多疾病联合建模,挖掘共享遗传机制

传统全基因组关联研究(GWAS)虽推动了复杂疾病理解,但常忽略非线性基因互作。深度学习可捕捉复杂基因模式,但现有方法多依赖特征选择,或受限于已知通路,或在全数据集应用时导致数据泄露。此外,协变量可能夸大预测性能,却不反映真实遗传信号。本文探索不同深度学习架构在GWAS中的表现,证明精心设计的结构可在严格无泄漏条件下超越现有方法。基于此,我们扩展为多标签框架,联合建模五种疾病,利用共享遗传结构提升效率与发现能力。在37,000例样本、500万SNP数据上,方法实现竞争性预测性能(AUC 0.68–0.96),提供一种可扩展、无泄漏且生物学意义明确的多病种GWAS分析新范式。

原文摘要 · Abstract (English)

Traditional GWAS has advanced our understanding of complex diseases but often misses nonlinear genetic interactions. Deep learning offers new opportunities to capture complex genomic patterns, yet existing methods mostly depend on feature selection strategies that either constrain analysis to known pathways or risk data leakage when applied across the full dataset. Further, covariates can inflate predictive performance without reflecting true genetic signals. We explore different deep learning architecture choices for GWAS and demonstrate that careful architectural choices can outperform existing methods under strict no-leakage conditions. Building on this, we extend our approach to a multi-label framework that jointly models five diseases, leveraging shared genetic architecture for improved efficiency and discovery. Applied to five million SNPs across 37,000 samples, our method achieves competitive predictive performance (AUC 0.68-0.96), offering a scalable, leakage-free, and biologically meaningful approach for multi-disease GWAS analysis.

GWAS深度学习多病建模遗传分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。