重复刺激导致脑电解码模型性能被高估4.5%-7.4%。
The Repeated-Stimulus Confound in Electroencephalography
- 用重复刺激训练和测试模型,导致刺激身份成为性能误判的干扰因素。
- 实际解码准确率被高估4.46%-7.42%,每1%误差对应0.26%过度估计。
- 该问题可被滥用以支持伪科学结论,如超感官认知存在。
在神经解码研究中,常使用参与者对刺激的反应记录来训练模型。近年来,深度学习成果大量应用于此类研究,但数据密集型模型推动了对更大数据集的需求。部分研究让同一刺激重复呈现多次以增加训练样本。然而,当模型在相同刺激上训练并评估时,刺激身份会成为准确率的混淆因子,我们称之为重复刺激混淆。我们识别出一个易受影响的数据集及16篇受此问题影响的论文。通过复现这些研究中的模型,发现其解码准确率被高估了4.46%-7.42%。分析表明,每1%的混淆导致的准确率提升,会使高估程度增加0.26%。该混淆不仅导致性能评估过于乐观,还削弱了相关论文中多个主张的有效性。我们进一步实验发现,相同方法可被用来支持一系列伪科学结论,如超感官知觉的存在。
原文摘要 · Abstract (English)
In neural-decoding studies, recordings of participants' responses to stimuli are used to train models. In recent years, there has been an explosion of publications detailing applications of innovations from deep-learning research to neural-decoding studies. The data-hungry models used in these experiments have resulted in a demand for increasingly large datasets. Consequently, in some studies, the same stimuli are presented multiple times to each participant to increase the number of trials available for use in model training. However, when a decoding model is trained and subsequently evaluated on responses to the same stimuli, stimulus identity becomes a confounder for accuracy. We term this the repeated-stimulus confound. We identify a susceptible dataset, and 16 publications which report model performance based on evaluation procedures affected by the confound. We conducted experiments using models from the affected studies to investigate the likely extent to which results in the literature have been misreported. Our findings suggest that the decoding accuracies of these models were overestimated by between 4.46-7.42%. Our analysis also indicates that per 1% increase in accuracy under the confound, the magnitude of the overestimation increases by 0.26%. The confound not only results in optimistic estimates of decoding performance, but undermines the validity of several claims made within the affected publications. We conducted further experiments to investigate the implications of the confound in alternative contexts. We found that the same methodology used within the affected studies could also be used to justify an array of pseudoscientific claims, such as the existence of extrasensory perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。