EEG情绪识别的准确率受评估流程影响,需区分不同实验设置下的结果意义。
Evaluation Protocols and Cross-Subject Generalization in EEG Emotion Recognition
- 分离评估流程中的目标、开发步骤和报告规则,揭示评估方式对结果的影响。
- 跨被试测试中,准确率从0.5348降至0.3954,显示模型泛化能力受限于个体差异。
- 建议按依赖被试、不相交被试、跨会话等不同问题分别报告结果,避免误读。
EEG情绪识别报告的准确率不仅取决于分类器,更受完整评估流程影响。本文将目标量、开发流程与报告规则分离,并以一个已存档的动力图卷积神经网络(DGCNN)路径在SEED和SEED-IV数据集上为例进行分析。在协议匹配的被试依赖测试中,SEED结果与公开参考值相差不超过1.47个百分点;而SEED-IV的3.40点差异仍未解决。在30条匹配的SEED被试-会话轨迹中,基于重复测试集评估选择检查点,使平均窗口准确率从第80轮的0.7855提升至0.8892。在五折被试互斥评估下,验证集选定的检查点在训练参与者上的试验准确率达0.9990(SEED)和0.9920(SEED-IV)。对于完全未见的被试,准确率为0.5348(95%条件被试级偏差校正加速[BCa]区间[0.4667, 0.5985]);SEED-IV估计为0.3954([0.3343, 0.4648]),仅作为次级敏感性证据报告,因其协议匹配兼容性检查未完成。观察到的训练到未见被试间的差距无法用简单优化欠拟合解释,但也无法排除被试身份、实现、预处理、表示或分布因素的影响。辅助分析表明,被试排名依赖于表示方式与时间尺度,而开发阶段选择的尾部风险集成在独立最终评估中未带来正收益。因此,应将被试依赖、被试互斥和跨会话结果视为回答不同问题。
原文摘要 · Abstract (English)
Reported accuracy in electroencephalography (EEG) emotion recognition depends on the complete evaluation procedure, not only the classifier. We separate the target quantity, development procedure, and reporting rule, then use one archived dynamical graph convolutional neural network (DGCNN) pathway on SEED and SEED-IV as an illustrative case. In a protocol-matched subject-dependent check, the SEED result was within 1.47 percentage points of the public reference value; the 3.40-point SEED-IV difference remained unresolved. Across 30 matched SEED subject-session trajectories, checkpoint selection based on repeated test-set evaluation increased mean window accuracy from 0.7855 at epoch 80 to 0.8892. Under five-fold subject-disjoint evaluation, validation-selected checkpoints achieved training-participant trial accuracies of 0.9990 on SEED and 0.9920 on SEED-IV. Accuracy for entirely held-out participants was 0.5348 (95% conditional subject-level bias-corrected and accelerated [BCa] interval [0.4667, 0.5985]) on SEED. The SEED-IV estimate was 0.3954 ([0.3343, 0.4648]) and is reported only as secondary sensitivity evidence because its protocol-matched compatibility check remained unresolved. The observed train-to-held-out-subject gaps are inconsistent with simple optimization underfitting, but they do not isolate subject identity from implementation, preprocessing, representation, or distributional factors. Supporting analyses further showed that participant rankings depended on representation and time scale, while a development-selected tail-risk ensemble did not establish a positive gain in a separate final evaluation. Subject-dependent, subject-disjoint, and cross-session results should therefore be reported as answers to different questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。