用不确定性估计提升阿尔茨海默病基因分类准确率
Uncertainty-Aware Genomic Classification of Alzheimer's Disease: A Transformer-Based Ensemble Approach with Monte Carlo Dropout
- 融合Transformer与随机森林,通过蒙特卡洛丢弃估算预测不确定性
- 剔除高不确定性样本后准确率提升10.24%,F1值提高23.62%
- 适合需要高可靠性基因分类的临床研究或辅助诊断场景
阿尔茨海默病(AD)遗传背景复杂,基于基因组数据的分类面临挑战。我们提出一种基于Transformer的集成模型TrUE-Net,结合全基因组测序(WGS)数据,利用蒙特卡洛丢弃进行不确定性估计。模型将保留单核苷酸多态性(SNP)序列结构的Transformer与使用展平基因型的随机森林并行结合。通过设定不确定性阈值,将样本分为高方差(不确定)和低方差(确定)两组。分析1050名个体,测试集占比一半。整体准确率为0.6514,受试者工作特征曲线下面积(AUC)为0.6636。剔除不确定样本后,准确率从0.6263升至0.7287(提升10.24%),F1值从0.5843增至0.8205(提升23.62%)。蒙特卡洛丢弃驱动的不确定性评估有助于识别需进一步临床评估的模糊病例,从而提升基因分类的可靠性。
原文摘要 · Abstract (English)
INTRODUCTION: Alzheimer's disease (AD) is genetically complex, complicating robust classification from genomic data. METHODS: We developed a transformer-based ensemble model (TrUE-Net) using Monte Carlo Dropout for uncertainty estimation in AD classification from whole-genome sequencing (WGS). We combined a transformer that preserves single-nucleotide polymorphism (SNP) sequence structure with a concurrent random forest using flattened genotypes. An uncertainty threshold separated samples into an uncertain (high-variance) group and a more certain (low-variance) group. RESULTS: We analyzed 1050 individuals, holding out half for testing. Overall accuracy and area under the receiver operating characteristic (ROC) curve (AUC) were 0.6514 and 0.6636, respectively. Excluding the uncertain group improved accuracy from 0.6263 to 0.7287 (10.24% increase) and F1 from 0.5843 to 0.8205 (23.62% increase). DISCUSSION: Monte Carlo Dropout-driven uncertainty helps identify ambiguous cases that may require further clinical evaluation, thus improving reliability in AD genomic classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。