整合多源遗传数据提升偏头痛预测准确率
EFGPP: Exploratory framework for genotype-phenotype prediction

- 构建可复现框架,融合基因、临床等多类数据
- 多源数据组合使预测AUC达0.688,优于单一数据
- 抑郁相关遗传信号对偏头痛有预测价值
从遗传数据预测复杂人类表型具有挑战性,因不同数据源仅包含部分信号。本文提出EFGPP框架,用于生成、排序和整合多种数据以进行基因型到表型的预测。在英国生物银行733名个体的偏头痛预测中,该框架结合了基因型特征、主成分、临床与代谢组协变量,以及来自偏头痛和抑郁症GWAS的多基因风险评分(使用PLINK、PRSice-2、AnnoPred、LDAK-GWAS生成)。单个数据类型最佳测试AUC为0.644,多数据融合后性能提升至0.688(偏头痛导向输入)和0.663(跨性状抑郁衍生输入)。基因型特征单独表现未超过仅协变量基线,但优于PRS单独使用,且抑郁来源的PRS展现出有效预测信号。EFGPP为复杂表型预测中异质遗传数据的优先级排序与整合提供了实用验证框架。
原文摘要 · Abstract (English)
Predicting complex human traits from genetic data is challenging because different genetic, clinical, and molecular data sources often contain different parts of the signal. Here, we present EFGPP, a reproducible framework for generating, ranking, and combining multiple types of data for genotype-to-phenotype prediction. We applied EFGPP to migraine prediction using UK Biobank data from 733 individuals. The framework combined genotype-derived features, principal components, clinical and metabolomic covariates, and polygenic risk scores generated from migraine and depression GWAS using PLINK, PRSice-2, AnnoPred, and LDAK-GWAS. The best single data type achieved a test AUC of 0.644, while combining multiple data types improved performance to 0.688 using migraine-focused inputs and 0.663 using cross-trait depression-derived inputs. Genetic features alone did not outperform the covariates-only baseline, but genotype-derived features performed better than PRS alone, and depression-derived PRS showed useful predictive signal. Overall, EFGPP provides a practical proof-of-concept framework for prioritising and integrating heterogeneous genetic data sources for complex phenotype prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。