87%的脑电情绪识别论文存在方法缺陷,导致结果虚高。
The Role of Review Process Failures in Affective State Estimation: An Empirical Investigation of DEAP Dataset
- 分析101篇论文,发现数据泄露、特征选择偏倚等常见问题
- 方法错误可使分类准确率虚高最高达46%
- 警示神经科学中机器学习研究的评审与标准亟待加强
基于脑电图(EEG)的情绪状态估计可靠性受到质疑,原因在于报告性能差异大且缺乏标准化评估协议。我们回顾了101项使用广泛DEAP数据集的研究,发现普遍存在方法学问题:数据分割不当导致数据泄露、特征选择存在偏倚、超参数优化不严谨、忽略类别不平衡以及方法描述不充分。值得注意的是,近87%的研究存在至少一项此类错误。通过实验分析,我们观察到这些方法学缺陷可使分类准确率虚高最高达46%。研究揭示了标准化评估实践的根本性缺失,并凸显了神经科学中机器学习应用在同行评审过程中的关键缺陷,强调亟需建立更严格的方法学标准与评估协议。
原文摘要 · Abstract (English)
The reliability of affective state estimation using EEG data is in question, given the variability in reported performance and the lack of standardized evaluation protocols. To investigate this, we reviewed 101 studies, focusing on the widely used DEAP dataset for emotion recognition. Our analysis revealed widespread methodological issues that include data leakage from improper segmentation, biased feature selection, flawed hyperparameter optimization, neglect of class imbalance, and insufficient methodological reporting. Notably, we found that nearly 87% of the reviewed papers contained one or more of these errors. Moreover, through experimental analysis, we observed that such methodological flaws can inflate the classification accuracy by up to 46%. These findings reveal fundamental gaps in standardized evaluation practices and highlight critical deficiencies in the peer review process for machine learning applications in neuroscience, emphasizing the urgent need for stricter methodological standards and evaluation protocols.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。