在严格控制条件下,脑电波解码发音元音的跨人可靠性证据有限。
Leakage-Audited Benchmarking Reveals Limited Evidence for Cross-Subject Auditory-Evoked EEG Vowel Perception Decoding

- 构建统一基准,严格控制实验与模型变量
- 随机森林最高准确率21.47%,仍低于随机水平
- 深度模型性能接近随机,结果高度依赖训练种子
我们在单一基准中控制试验身份、模型身份、预测来源和个体推断,测试了听觉诱发脑电图是否支持跨被试的五元音感知解码。从OpenNeuro ds006104 v1.0.1重建研究2事件表,分析辅音-元音对任务。一对一标记-刺激配对生成3,840个独立试验;经对照条件选择与伪迹剔除后保留1,094个时间段,来自16名参与者和61个脑电通道。使用留一被试法评估13种不同实现方式,基于36,102次试验预测重构被试级指标,共33次完整预测副本。随机森林在平衡准确率上表现最佳(21.474%,95%置信区间19.526%-23.482%;随机水平为20%),但其被试级检验及所有实现均未通过校正。深度模型性能接近随机,多个架构表现出显著的种子依赖性与低试验标签一致性。探索性多数据集建模分析显示,9,616次真实重拟合在3-15人训练组中无单调性能提升。在此数据集与协议下,跨被试五元音解码的可靠证据有限。该基准提供从原始数据到保留时段、预测、被试级指标、多重性校正推断与边界诊断分析的可复现链条。
原文摘要 · Abstract (English)
We tested whether auditory-evoked EEG supports subject-independent five-vowel perception decoding when trial identity, model identity, prediction provenance, and participant-level inference are controlled within a single benchmark. We reconstructed Study 2 event tables from OpenNeuro ds006104 version 1.0.1 and analysed the consonant-vowel pair task. One-to-one marker-stimulus pairing yielded 3,840 independent trials; control-condition selection and artifact rejection retained 1,094 epochs from 16 participants and 61 EEG channels. Thirteen unique implementations were evaluated using leave-one-subject-out testing, with participant metrics reconstructed from 36,102 trial predictions across 33 complete prediction replicas. Random Forest was numerically highest at 21.474% balanced accuracy (95% participant-bootstrap interval, 19.526-23.482%; chance, 20%), but neither its participant-level tests nor any implementation survived correction across the 13-model family. Deep-model performance was close to chance, and several architectures showed substantial seed-dependent variation and low trial-label agreement. An exploratory MDM analysis comprising 9,616 genuine refits across training cohorts of 3-15 participants showed no monotonic performance gain. Within this dataset and protocol, evidence for reliable cross-subject five-vowel decoding is limited. The benchmark provides a reproducible chain from source rows to retained epochs, predictions, participant-level metrics, multiplicity-adjusted inference, and bounded diagnostic analyses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。