研究脑电图深度学习中数据划分对模型性能的影响,避免结果夸大。
The role of data partitioning on the performance of EEG-based deep learning models in supervised cross-subject analysis: a preliminary study
- 采用基于被试的交叉验证,防止数据泄露和过拟合。
- 嵌套交叉验证比非嵌套更可靠,能提升模型泛化能力。
- 适用于脑电信号分类研究者,尤其关注临床疾病检测的团队。
深度学习正显著推动脑电图(EEG)数据分析,能有效挖掘信号中的高度非线性模式。数据划分与交叉验证对评估模型性能及确保研究可比性至关重要,因其可能因信号特性(如生物特征)导致结果差异和数据泄露,进而引发研究不可比、性能被高估等问题。然而,该领域尚无明确的数据划分与交叉验证指南,也缺乏对其影响的量化评估。为帮助研究者制定最优实验策略,本文系统比较了五种交叉验证设置在三个监督式跨被试分类任务(运动想象、帕金森病、阿尔茨海默病检测)上的表现,覆盖四种复杂度递增的模型架构(ShallowConvNet、EEGNet、DeepConvNet、Temporal-based ResNet)。通过对超过10万次模型训练的分析表明:首先,除可接受组内分析的情况(如运动想象),应优先使用基于被试的交叉验证;其次,嵌套方法(N-LNSO)比非嵌套方法更可靠,后者易受数据泄露影响,偏好大模型在验证集上过拟合。本研究为EEG深度学习研究提供了数据划分与交叉验证的分析依据,并提出避免数据泄露的实践建议,以遏制当前领域普遍存在的性能高估现象。
原文摘要 · Abstract (English)
Deep learning is significantly advancing the analysis of electroencephalography (EEG) data by effectively discovering highly nonlinear patterns within the signals. Data partitioning and cross-validation are crucial for assessing model performance and ensuring study comparability, as they can produce varied results and data leakage due to specific signal properties (e.g., biometric). Such variability leads to incomparable studies and, increasingly, overestimated performance claims, which are detrimental to the field. Nevertheless, no comprehensive guidelines for proper data partitioning and cross-validation exist in the domain, nor is there a quantitative evaluation of their impact on model accuracy, reliability, and generalizability. To assist researchers in identifying optimal experimental strategies, this paper thoroughly investigates the role of data partitioning and cross-validation in evaluating EEG deep learning models. Five cross-validation settings are compared across three supervised cross-subject classification tasks (BCI, Parkinson's, and Alzheimer's disease detection) and four established architectures of increasing complexity (ShallowConvNet, EEGNet, DeepConvNet, and Temporal-based ResNet). The comparison of over 100,000 trained models underscores, first, the importance of using subject-based cross-validation strategies for evaluating EEG deep learning models, except when within-subject analyses are acceptable (e.g., BCI). Second, it highlights the greater reliability of nested approaches (N-LNSO) compared to non-nested counterparts, which are prone to data leakage and favor larger models overfitting to validation data. In conclusion, this work provides EEG deep learning researchers with an analysis of data partitioning and cross-validation and offers guidelines to avoid data leakage, currently undermining the domain with potentially overestimated performance claims.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。