针对脑电运动想象解码的个体差异,构建大规模基准并提出轻量级个性化方案。
Subject-Level Heterogeneity in EEG Motor Imagery Decoding: A Large-Scale Benchmark and Portfolio-Based Reduction of the Search Space

- 在三大公开数据集上系统评测超20万种解码流程,揭示个体差异显著。
- 采用组合策略的精简管道包可保留94%以上最优性能,实现高效个性化。
- 发现不同人适合不同算法,为自适应脑机接口设计提供实证支持。
稳健的脑电运动想象解码受限于强烈的个体间差异,难以找到跨用户通用的解码流程。本文在三个公开数据集(Cho2017:52人,PhysionetMI:109人,Zhou2016:4人)上构建了大规模、标准化的会话内基准,使用统一的MOABB LeftRightImagery设置,两个频段(8-15 Hz 和 8-30 Hz),以及广泛的特征提取、预处理与分类组合,分析了216,714条原始评估结果,经结构化聚合后得到44,928、109,000和4,192个受试者级观测值。协方差切线空间投影(cov-tgsp)和共空间模式(CSP)始终表现最强方法族,但其相对排序依赖数据集。在Cho2017中,8-30 Hz下cov-tgsp平均准确率达0.712±0.140;在Zhou2016中,8-15 Hz下CSP达0.832±0.121。这些总体排名掩盖了严重的受试者级异质性:Cho2017中有42种胜出流程,PhysionetMI中有93种。随后利用该基准构建大小为K的解码流程组合包,比较了基于排名的Top-K Mean等方法。结果一致表明,Top-K Mean表现最佳。单个最优全局流程已保留Cho2017中94.2%和PhysionetMI中81.8%的基准性能;当K=12时,性能提升至96.5%和90.0%。因此,解码性能高度依赖个体,可通过紧凑组合包有效利用异质性实现个性化。
原文摘要 · Abstract (English)
Robust EEG motor imagery decoding remains limited by strong inter-individual variability, making it difficult to identify pipelines that generalize across users. We present a large-scale, standardized within-session benchmark of decoding pipelines across three public datasets: Cho2017 (52 subjects), PhysionetMI (109 subjects), and Zhou2016 (4 subjects). Using a common MOABB LeftRightImagery setting, two frequency bands (8-15 Hz and 8-30 Hz), and a broad combination of feature extraction, preprocessing, and classification steps, we analyzed 216,714 raw evaluation rows, which after structured aggregation yielded 44,928, 109,000, and 4,192 subject-level observations respectively. Covariance tangent-space projection (cov-tgsp) and Common Spatial Patterns (CSP) consistently defined the strongest methodological families, though their relative ordering was dataset-dependent. On Cho2017, the best family-level mean accuracy came from cov-tgsp in 8-30 Hz (0.712 +/- 0.140), whereas Zhou2016 favored CSP (0.832 +/- 0.121 in 8-15 Hz). These aggregate rankings concealed substantial subject-level heterogeneity: 42 distinct winning pipelines across 52 Cho2017 subjects, and 93 across 109 PhysionetMI subjects. We then used the benchmark as an empirical performance landscape for building compact portfolios of pipelines of size K. Several construction procedures were compared, including a ranking-based Top-K Mean heuristic and search-based strategies. Results were broadly consistent, with Top-K Mean giving the best trade-off. A single best global pipeline already retained 94.2% of the oracle in Cho2017 and 81.8% in PhysionetMI; at K = 12, oracle retention rose to 96.5% and 90.0%. The landscape is therefore subject-dependent, and this heterogeneity can be exploited through compact portfolios that make personalization more feasible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。