CLEAN-MI提升运动想象脑机接口数据质量与模型性能
CLEAN-MI: A Scalable and Efficient Pipeline for Constructing High-Quality Neurodata in Motor Imagery Paradigm
- 通过频带滤波、电极模板选择等多步骤清洗,系统化处理多源脑电数据
- 在多个公开数据集上显著提升数据质量和分类准确率
- 适合需要高质量脑电数据的脑机接口研究者使用
构建大规模、高质量的数据集是发展稳健且可泛化的运动想象(MI)脑-机接口(BCI)基础模型的关键前提。然而,来自不同受试者和设备的脑电(EEG)信号常面临信噪比低、电极配置异质性及显著个体差异等问题,严重制约模型训练效果。本文提出CLEAN-MI,一个可扩展、系统化的数据构建流程,用于构建大规模、高效且精准的运动想象范式神经数据。该方法融合频带滤波、通道模板选择、受试者筛选与边缘分布对齐,系统剔除无关或低质量数据,并标准化多源EEG数据集。我们在多个公开的MI数据集上验证了CLEAN-MI的有效性,均实现数据质量与分类性能的持续提升。
原文摘要 · Abstract (English)
The construction of large-scale, high-quality datasets is a fundamental prerequisite for developing robust and generalizable foundation models in motor imagery (MI)-based brain-computer interfaces (BCIs). However, EEG signals collected from different subjects and devices are often plagued by low signal-to-noise ratio, heterogeneity in electrode configurations, and substantial inter-subject variability, posing significant challenges for effective model training. In this paper, we propose CLEAN-MI, a scalable and systematic data construction pipeline for constructing large-scale, efficient, and accurate neurodata in the MI paradigm. CLEAN-MI integrates frequency band filtering, channel template selection, subject screening, and marginal distribution alignment to systematically filter out irrelevant or low-quality data and standardize multi-source EEG datasets. We demonstrate the effectiveness of CLEAN-MI on multiple public MI datasets, achieving consistent improvements in data quality and classification performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。