arXiv:2605.08495cs.LGq-bio.NC2026-05被引 9

构建统一框架评测脑电模型,发现大模型优势有限且多数任务仍难攻克。

NeuralBench: A Unifying Framework to Benchmark NeuroAI Models

论文配图:NeuralBench: A Unifying Framework to Benchmark NeuroAI Models
图 1 · 摘自论文原文
  • 提出NeuralBench统一框架,标准化脑电模型评估流程
  • 包含36项任务、14种模型,在94个数据集上验证效果
  • 适合脑机接口、临床诊断研究者参考,推动评测标准化

深度学习与大型公开数据集推动了脑电记录处理AI模型的快速发展。然而,系统性评估这些模型仍面临挑战:预处理、训练方式差异大,下游评估常局限于少量任务或数据集。为此,我们推出NeuralBench:一个统一的脑活动AI模型基准框架。配套发布NeuralBench-EEG v1.0,包含36项脑电(EEG)任务、14种深度学习架构,并在94个通过标准化接口访问的数据集上进行评估。该首次发布的脑电专用版本揭示两个关键发现:当前基础模型仅略优于特定任务模型;许多任务(如认知解码、临床预测)对最佳模型仍具挑战性。NeuralBench设计支持新任务、数据集、模型及神经影像模态的集成,已初步扩展至MEG和fMRI。本文邀请社区共同建设这一开源框架,推动神经影像模型的统一基准标准。

原文摘要 · Abstract (English)

Deep learning and large public datasets have recently catalyzed the proliferation of AI models for processing brain recordings. However, systematically evaluating these models remains a challenge: not only do the preprocessing pipelines, training and finetuning approaches largely vary across studies, but their downstream evaluation is often limited to small sets of tasks and/or datasets. Here, we present NeuralBench: a unified framework for benchmarking AI models of brain activity. We accompany this framework with NeuralBench-EEG v1.0 -- a large EEG benchmark that includes 36 electroencephalography (EEG) tasks and 14 deep learning architectures, and is evaluated on 94 datasets accessed through a standardized interface. This first EEG-focused release already highlights two main findings. First, current foundation models only marginally outperform task-specific models. Second, a large set of tasks (e.g. cognitive decoding, clinical predictions) remain highly challenging, even for the best models. Critically, NeuralBench is designed for the integration of new tasks, datasets, models, and neuroimaging modalities, as illustrated by preliminary extensions to MEG and fMRI datasets and models. Through this white paper, we invite the community to expand this open-source framework and work together toward a unified benchmarking standard for neuroimaging models.

脑电建模模型评估神经人工智能基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。