提出临床脑电图模型评估新协议,防止数据集偏差干扰结果
A Negative-Control Protocol for Clinical EEG Foundation-Model Benchmarks: Dataset Identity and External-Cohort Stress Testing

- 用不重叠受试者验证,检测模型是否依赖特定数据集
- 经典特征在认知障碍分类中表现最佳(0.734宏AUROC)
- 适合关注临床脑电模型可靠性的研究者参考
脑电图基础模型的性能可能受队列、电极布局或探测设计影响。我们在四个基准数据集及韩国CAUEEG上评估了五种模型在五个任务中的表现,采用受试者无交集的验证策略。CAUEEG为记录级数据,包含标注的非重叠保留敏感性。在匹配的正常/轻度认知障碍/痴呆分类任务中(1,187条记录),经典特征达到0.734宏AUROC(灵敏度0.736),优于BIOT-bipolar16的0.677、CBraMod的0.669和REVE的0.568。标注的非重叠子集仍保持经典特征优于REVE的排序(0.717对0.565)。所有五种编码器在主成分分析前后的数据集身份识别准确率均为1.000;标签打乱后降至随机水平,而平衡采样仍保持1.000。这表明模型性能由数据集归属决定,而非因果地点、地理或人群效应。一个完全随机初始化的编码器在CAUEEG上表现优于预训练的REVE(0.667对0.570),且正确与打乱源标签的LoRA运行在非匹配描述性敏感性上数值相近,无法得出标签效应或等效性结论。在CHB-MIT跨受试者发作检测中,REVE达到0.793 AUROC,优于最佳对比方法的0.739,差值为+5.38个百分点(95% CI -0.36至+11.22),故对比方法优劣未定;且幅度感知对比方法无法从保留的归一化输入中重建。本文提炼出电极布局匹配、患者重叠检查、更强对比模型和表示控制等要素,形成临床脑电图基础模型研究的报告规范。
原文摘要 · Abstract (English)
EEG foundation-model gains may depend on cohort, montage, or probe design. We evaluated five models on five tasks across four benchmark datasets plus Korean CAUEEG, using subject-disjoint validation where identifiers exist. CAUEEG is recording-level with an annotated no-overlap held-out sensitivity. On matched CAUEEG normal/mild cognitive impairment/dementia classification (1,187 recordings), classical features reached 0.734 macro-AUROC (enhanced sensitivity: 0.736), versus BIOT-bipolar16 0.677, CBraMod 0.669, and REVE 0.568. The annotated no-overlap held-out subset preserved the classical-over-REVE ordering (0.717 versus 0.565). All five encoders decoded dataset identity at 1.000 before and after in-fold PCA-50; label permutations collapsed to chance and balanced subsamples remained at 1.000. This establishes dataset membership, not a causal site, geography, or population effect. A matched fully randomly initialized encoder was descriptively higher than pretrained REVE on CAUEEG (0.667 versus 0.570), and correct- versus scrambled-source-label LoRA runs yielded numerically similar AUROCs in unmatched descriptive sensitivities, not label-effect estimates or equivalence tests. On CHB-MIT cross-subject ictal detection, REVE reached 0.793 AUROC, versus 0.739 for the best tested enhanced nonlinear comparator, 0.691 for fully random initialization, and 0.505 for raw-signal random features. The paired REVE-minus-enhanced-comparator difference was +5.38 percentage points (95% CI -0.36 to +11.22), so comparator superiority remains unresolved; an amplitude-aware comparator also cannot be reconstructed from the retained normalized inputs. We distill montage matching, patient-overlap checks, stronger comparators, and representation controls into a reporting protocol for clinical EEG foundation-model studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。