通过数据融合与匹配,用小样本实现高精度阿尔茨海默病风险预测。
Comprehensive Methodology for Sample Augmentation in EEG Biomarker Studies for Alzheimers Risk Classification
- 整合多源脑电数据,用倾向评分匹配平衡人群差异。
- 在2:1到10:1匹配比下,分类准确率达0.92至0.96。
- 适合神经退行性疾病早期筛查研究者参考。
痴呆症是全球健康挑战,阿尔茨海默病(AD)占约70%病例。脑电图(EEG)在识别AD风险方面有潜力,但大规模样本获取困难。本研究融合信号处理、数据调和与统计方法,提升样本量并增强风险分类可靠性。利用四个数据库的脑电数据,通过调和处理消除站点效应,同时保留年龄、性别等协变量。采用倾向评分匹配(PSM)设定2:1、5:1、10:1三种比例,评估样本量对模型性能的影响。最终数据集经决策树与交叉验证进行机器学习分析。结果表明,通过PSM平衡样本显著提升分类准确率,范围在0.92至0.96之间。该方法即使在样本有限时也能实现精准风险识别,为其他神经退行性疾病研究提供可行路径。
原文摘要 · Abstract (English)
Background: Dementia, marked by cognitive decline, is a global health challenge. Alzheimer's disease (AD), the leading type, accounts for ~70% of cases. Electroencephalography (EEG) measures show promise in identifying AD risk, but obtaining large samples for reliable comparisons is challenging. Objective: This study integrates signal processing, harmonization, and statistical techniques to enhance sample size and improve AD risk classification reliability. Methods: We used advanced EEG preprocessing, feature extraction, harmonization, and propensity score matching (PSM) to balance healthy non-carriers (HC) and asymptomatic E280A mutation carriers (ACr). Data from four databases were harmonized to adjust site effects while preserving covariates like age and sex. PSM ratios (2:1, 5:1, 10:1) were applied to assess sample size impact on model performance. The final dataset underwent machine learning analysis with decision trees and cross-validation for robust results. Results: Balancing sample sizes via PSM significantly improved classification accuracy, ranging from 0.92 to 0.96 across ratios. This approach enabled precise risk identification even with limited samples. Conclusion: Integrating data processing, harmonization, and balancing techniques improves AD risk classification accuracy, offering potential for other neurodegenerative diseases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。