用Transformer模型分析基因数据,精准预测肺癌风险并解释关键基因位点。
HEMERA: A Human-Explainable Transformer Model for Estimating Lung Cancer Risk using GWAS Data
- 直接处理原始基因型数据,通过位置编码和嵌入学习提升特征表达。
- 在27,254人数据上达到99%以上AUC,性能优于传统方法。
- 可解释性强,能定位与肺癌相关的关键基因位点,适合临床研究使用。
肺癌是美国第三大常见癌症,也是癌症死亡的首要原因。尽管吸烟是主要风险因素,但非吸烟者中肺癌的发生及家族聚集性研究提示存在遗传成分。通过全基因组关联研究(GWAS)识别的遗传生物标志物是评估肺癌风险的有力工具。本文提出HEMERA(基于GWAS数据的人类可解释变压器模型,用于估计肺癌风险),一种采用可解释的变压器深度学习框架,直接处理单核苷酸多态性(SNPs)的原始基因型数据以预测肺癌风险。与以往方法不同,HEMERA无需临床协变量,引入加性位置编码、神经基因型嵌入和优化变异过滤策略。基于层间积分梯度的后处理可解释模块可将模型预测归因于特定SNPs,结果与已知肺癌风险位点高度一致。模型在27,254名退伍军人计划参与者的数据上训练,获得超过99%的受试者工作特征曲线下面积(AUC)得分。这些发现支持了透明化、假设生成型的个性化肺癌风险评估模型,为早期干预提供可能。
原文摘要 · Abstract (English)
Lung cancer (LC) is the third most common cancer and the leading cause of cancer deaths in the US. Although smoking is the primary risk factor, the occurrence of LC in never-smokers and familial aggregation studies highlight a genetic component. Genetic biomarkers identified through genome-wide association studies (GWAS) are promising tools for assessing LC risk. We introduce HEMERA (Human-Explainable Transformer Model for Estimating Lung Cancer Risk using GWAS Data), a new framework that applies explainable transformer-based deep learning to GWAS data of single nucleotide polymorphisms (SNPs) for predicting LC risk. Unlike prior approaches, HEMERA directly processes raw genotype data without clinical covariates, introducing additive positional encodings, neural genotype embeddings, and refined variant filtering. A post hoc explainability module based on Layer-wise Integrated Gradients enables attribution of model predictions to specific SNPs, aligning strongly with known LC risk loci. Trained on data from 27,254 Million Veteran Program participants, HEMERA achieved >99% AUC (area under receiver characteristics) score. These findings support transparent, hypothesis-generating models for personalized LC risk assessment and early intervention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。