arXiv:2508.17742eess.SPcs.AI2025-08中稿 · ICML被引 17

首个标准化脑电基础模型评测基准,解决评估不统一问题。

EEG-FM-Bench: A Comprehensive Benchmark for the Systematic Evaluation and Diagnostic Analyses of EEG Foundation Models

  • 构建统一评测体系,整合14个数据集和多种实验设置。
  • 发现多任务学习可缓解数据少时过拟合,但存在负迁移风险。
  • 模型性能受目标对齐与脑电设计影响,超越单纯规模比拼。

脑电图基础模型(EEG-FMs)推动了脑信号分析的发展,但缺乏标准化评估基准阻碍了模型比较与科学进展。现有评估采用不一致的协议,导致跨模型比较不可靠,且缺乏诊断分析,难以揭示迁移效率与缩放行为的内在机制。为此,我们提出 extbf{EEG-FM-Bench},一个用于EEG-FMs标准化评估的统一系统。该基准整合了14个数据集,覆盖10种范式,包含多种微调策略、任务组织方式及分类器配置,并配备梯度与表征分析工具。实验与分析揭示:(1)多任务学习在数据稀缺的脑电场景中常作为有效正则化手段,减少过拟合,但在特定任务范式下可能引发负迁移;(2)预训练效率受限于重建目标与下游任务间的梯度冲突;(3)在已发布检查点与匹配下游协议下,模型或数据规模无法完全解释迁移性能,目标对齐、适应兼容性与脑电特性设计更为关键。该基准支持公平比较与可复现分析,推动更公正、可解释的EEG-FM研究。代码见https://github.com/xw1216/EEG-FM-Bench。

原文摘要 · Abstract (English)

Electroencephalography foundation models (EEG-FMs) have advanced brain signal analysis, but the lack of standardized evaluation benchmarks impedes model comparison and scientific progress. Current evaluations rely on inconsistent protocols that render cross-model comparisons unreliable, while a lack of diagnostic analyses obscures the internal mechanisms driving transfer efficiency and scaling behaviors. To address this, we introduce \textbf{EEG-FM-Bench}, a unified system for the standardized evaluation of EEG-FMs. The benchmark integrates 14 datasets across 10 paradigms and incorporates diverse experimental settings, including multiple fine-tuning strategies, task organizations, and classifier configurations, supported by tools for gradient and representation analysis. Our experiments and analysis reveal several critical insights: (1) multi-task learning often acts as a useful regularizer that mitigates overfitting in data-scarce EEG contexts, although negative transfer can arise under specific task paradigms; (2) pre-training efficiency is currently limited by gradient conflicts between reconstruction objectives and downstream tasks; (3) under released checkpoints and a matched downstream protocol, model or data scale alone does not fully explain transfer performance, while objective alignment, adaptation compatibility, and EEG-specific design appear to be important factors. This benchmark enables fair comparison and reproducible analysis, providing a step toward fairer comparison and more interpretable analysis of EEG-FMs. Code is available at https://github.com/xw1216/EEG-FM-Bench.

脑电图基础模型评测基准可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。