arXiv:2605.14698cs.LGcs.AI2026-05被引 3

构建最大规模脑电基准测试,评估大模型在临床与脑机接口中的实际表现。

NeuroAtlas: Benchmarking Foundation Models for Clinical EEG and Brain-Computer Interfaces

论文配图:NeuroAtlas: Benchmarking Foundation Models for Clinical EEG and Brain-Computer Interfaces
图 1 · 摘自论文原文
  • 整合42个数据集、26万小时脑电数据,覆盖癫痫、睡眠和脑龄等临床任务。
  • 发现通用时间序列模型性能不输专用脑电模型,且现有指标难以反映临床价值。
  • 提出事件级决策、睡眠分期特征等新评估方式,适合临床研究者参考。

基础模型(FMs)有望提取可跨下游任务泛化的统一表征,已在多个领域兴起,包括脑电图(EEG),但其在该领域的有效性尚不明确。已有评估在数据集、预处理方式及评价指标上存在差异,常掩盖临床相关性。我们推出NeuroAtlas,目前最大的脑电基准:涵盖42个数据集、260,000小时的临床脑电(癫痫、睡眠医学、脑龄估计)与脑机接口数据,并包含每项任务的多个数据集及定制化临床评估指标。除对比监督基线外,还评估了通用时间序列基础模型。结果有三:第一,专用于脑电的基础模型并未持续优于未针对脑电设计、也未在脑电上预训练的时间序列模型;第二,标准机器学习指标不足以评估临床实用性,因此我们深入评估了事件级决策质量、基于睡眠分期的特征表现以及癫痫、睡眠和脑龄领域的脑龄差距等更合适的指标;第三,同一领域内模型排名与性能差异显著。结论是,当前预训练模型整体表现相近,仅少数略优,尚未实现‘开箱即用’的统一脑电模型承诺。NeuroAtlas揭示了这一差距,并为下一代统一脑电基础模型提供了数据与评估工具。

原文摘要 · Abstract (English)

Foundation models (FMs) promise to extract unified representations that generalize across downstream tasks. They have emerged across fields, including electroencephalography (EEG), but it is less clear how effective they are in this particular field. Published evaluations differ in datasets, in the EEG-specific preprocessing that might influence reported results, and in the reported metrics, frequently obscuring the clinical relevance in EEG. We introduce NeuroAtlas, the largest EEG benchmark to date: 42 datasets and 260k hours covering clinical EEG (epilepsy, sleep medicine, brain age estimation) and brain-computer interfaces, and include multiple datasets per task along with bespoke clinical evaluation metrics. Besides evaluating EEG-FMs with respect to supervised baselines, we present results from generic time-series FMs. We report three findings. First, EEG-specific FMs do not consistently outperform time-series FMs, which have neither EEG-focused architectures nor been pretrained on EEG. Second, standard machine learning metrics are insufficient to assess clinical utility: thus, we thoroughly evaluate more appropriate measures such as the quality of event-level decision-making, hypnogram-derived features, and the brain-age gap in the domains of epilepsy, sleep, and brain age, respectively. Third, model rankings and performance can vary substantially within domains. We conclude that pretrained models perform largely on par, with only narrow advantages for a few, and that current models do not yet deliver on the promise of an out-of-the-box unified EEG model. NeuroAtlas exposes this gap and provides the datasets and metrics for the next generation of unified EEG FMs.

脑电图基础模型临床评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。