arXiv:2503.13497eess.SPcs.LG2025-03NeurIPS被引 8

研究发现脑电数据参与者多样性不足会严重限制模型泛化能力。

Is Limited Participant Diversity Impeding EEG-based Machine Learning?

  • 通过多层次数据生成框架,分析样本量与参与者多样性对模型的影响
  • 参与者分布差异导致性能提升受限,即使样本量增加也难突破瓶颈
  • 为脑电机器学习的数据收集和算法设计提供可操作建议

将脑电信号应用到机器学习具有推动神经科学研究和临床应用的巨大潜力。然而,基于脑电的机器学习模型的泛化性和鲁棒性通常依赖于训练数据的数量与多样性。目前普遍做法是将脑电记录切分为小段,从而显著增加样本数量,远超参与者的数量。本文将此视为多层级数据生成过程,通过大规模实证研究,考察模型性能随总体样本量和参与者多样性变化的缩放规律,并利用相同框架评估数据增强与自监督学习等应对数据有限问题的策略有效性。研究发现,模型性能的提升可能因参与者分布偏移而严重受限,并为数据收集与机器学习研究提供了可操作指导。实验代码已公开。

原文摘要 · Abstract (English)

The application of machine learning (ML) to electroencephalography (EEG) has great potential to advance both neuroscientific research and clinical applications. However, the generalisability and robustness of EEG-based ML models often hinge on the amount and diversity of training data. It is common practice to split EEG recordings into small segments, thereby increasing the number of samples substantially compared to the number of individual recordings or participants. We conceptualise this as a multi-level data generation process and investigate the scaling behaviour of model performance with respect to the overall sample size and the participant diversity through large-scale empirical studies. We then use the same framework to investigate the effectiveness of different ML strategies designed to address limited data problems: data augmentations and self-supervised learning. Our findings show that model performance scaling can be severely constrained by participant distribution shifts and provide actionable guidance for data collection and ML research. The code for our experiments is publicly available online.

脑电机器学习数据多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。