arXiv:2511.01951cs.LG2025-11

神经信号去噪新方法,自动清理脑电数据中的干扰。

NeuroClean: A Generalized Machine-Learning Approach to Neural Time-Series Conditioning

  • 无监督多步骤流程,自动识别并去除噪声成分。
  • 处理后模型分类准确率达97%,显著优于原始数据的74%。
  • 适合需要高可靠性的脑电信号分析研究者使用。

脑电图(EEG)和局部场电位(LFP)是记录大脑电活动的常用技术,广泛应用于临床与科研。然而,这些信号常受多种非脑源性伪迹和噪声干扰。为实现大规模数据的全自动预处理,本文提出NeuroClean管道,一种无监督、多功能的EEG/LFP预处理方法。该流程包含五步:带通滤波、工频噪声抑制、坏通道剔除,以及基于聚类算法的自动独立成分分析(ICA)成分筛选。通过机器学习分类器确保任务相关信号在清洗过程中被保留。在多个数据集上验证表明,NeuroClean可有效去除多种常见伪迹。在不同复杂度运动任务中,经清洗后的数据使优化的多项式逻辑回归模型分类准确率达到97%(随机水平为33.3%),远高于原始数据的74%。结果证明NeuroClean具有良好的泛化能力,适用于未来机器学习分析工作流。

原文摘要 · Abstract (English)

Electroencephalography (EEG) and local field potentials (LFP) are two widely used techniques to record electrical activity from the brain. These signals are used in both the clinical and research domains for multiple applications. However, most brain data recordings suffer from a myriad of artifacts and noise sources other than the brain itself. Thus, a major requirement for their use is proper and, given current volumes of data, a fully automatized conditioning. As a means to this end, here we introduce an unsupervised, multipurpose EEG/LFP preprocessing method, the NeuroClean pipeline. In addition to its completeness and reliability, NeuroClean is an unsupervised series of algorithms intended to mitigate reproducibility issues and biases caused by human intervention. The pipeline is designed as a five-step process, including the common bandpass and line noise filtering, and bad channel rejection. However, it incorporates an efficient independent component analysis with an automatic component rejection based on a clustering algorithm. This machine learning classifier is used to ensure that task-relevant information is preserved after each step of the cleaning process. We used several data sets to validate the pipeline. NeuroClean removed several common types of artifacts from the signal. Moreover, in the context of motor tasks of varying complexity, it yielded more than 97% accuracy (vs. a chance-level of 33.3%) in an optimized Multinomial Logistic Regression model after cleaning the data, compared to the raw data, which performed at 74% accuracy. These results show that NeuroClean is a promising pipeline and workflow that can be applied to future work and studies to achieve better generalization and performance on machine learning pipelines.

脑电分析信号去噪无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。