arXiv:2411.10450eess.SPcs.AI2024-11

清理脑电数据中的噪声,提升模型泛化能力。

Dataset Refinement for Improving the Generalization Ability of the EEG Decoding Model

  • 基于训练影响度量自动识别并剔除噪声数据。
  • 在两个公开数据集上验证,模型泛化性能显著提升。
  • 适合做脑机接口深度学习研究的团队参考。

脑电图(EEG)因其非侵入性和便捷性,被广泛应用于脑机接口中,是理解人类意图的有效工具。近年来,研究者们利用深度学习方法从EEG信号中解码人类意图。然而,由于EEG信号在采集过程中极易受噪声干扰,数据集中存在噪声数据的可能性很高。尽管早期研究通常假设数据集已充分清洗,但这一假设在实际的EEG数据集中并不总成立。本文提出一种数据集精炼算法,通过评估数据在训练过程中的影响程度来识别并剔除噪声数据。我们将该算法应用于两个运动想象类EEG公开数据集,并在三种不同模型上进行数据精炼与重训练。结果表明,使用精炼后的数据集重新训练模型,其泛化性能均优于原始数据集。因此,我们证明仅通过去除训练数据中的噪声,即可有效提升深度学习模型在EEG领域的泛化能力。

原文摘要 · Abstract (English)

Electroencephalography (EEG) is a generally used neuroimaging approach in brain-computer interfaces due to its non-invasive characteristics and convenience, making it an effective tool for understanding human intentions. Therefore, recent research has focused on decoding human intentions from EEG signals utilizing deep learning methods. However, since EEG signals are highly susceptible to noise during acquisition, there is a high possibility of the existence of noisy data in the dataset. Although pioneer studies have generally assumed that the dataset is well-curated, this assumption is not always met in the EEG dataset. In this paper, we addressed this issue by designing a dataset refinement algorithm that can eliminate noisy data based on metrics evaluating data influence during the training process. We applied the proposed algorithm to two motor imagery EEG public datasets and three different models to perform dataset refinement. The results indicated that retraining the model with the refined dataset consistently led to better generalization performance compared to using the original dataset. Hence, we demonstrated that removing noisy data from the training dataset alone can effectively improve the generalization performance of deep learning models in the EEG domain.

脑电图数据清洗泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。