arXiv:2608.04156cs.AIcs.LG2026-08

评测大模型对脑电图的综合理解能力,涵盖从指令到报告的全流程分析。

BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding

论文配图:BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding
图 1 · 摘自论文原文
  • 构建统一基准,支持指令驱动的脑电图全流程分析。
  • 覆盖17个数据集、超10万次执行,验证多模型在不同任务中的表现差异。
  • 适合研究脑机接口、医疗AI或大模型跨模态理解的开发者使用。

脑电图(EEG)分析不仅限于标签化,还需结合自然语言指令、信号处理、量化证据与科学解释。我们称此能力为「综合脑电理解」。现有评估多聚焦孤立解码任务或特定系统演示,难以全面衡量大语言模型(LLM)的性能。本文提出 enchmarkname{},一个统一的综合指令驱动脑电理解基准。包含四大子集:基础分析、睡眠评估、神经认知评估和生理整合,覆盖17个数据集、 umcases{} 项任务和超过 uminstances{} 个真实数据实例。给定指令及可选生理信号,系统需完成分析并生成科学报告,必要时输出伪影。评估涵盖数值、分类、集合、序列、语义及伪影验证。我们在两种范式下评测 ummodels{} 个代表性模型,执行超10万次:基于CodeAct的自主代码执行,以及基于BrainAgent的结构化智能体分析。结果表明,模型表现随任务、难度与执行方式显著变化,说明脑电理解能力依赖于模型本身及其运作方式。enchmarkname{} 提供可复现的测试平台,推动基于大模型的脑电理解发展。代码与基准将很快发布,评估结果将持续更新。

原文摘要 · Abstract (English)

Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural-language instructions, signal processing, quantitative evidence, and scientific interpretation. We term this capability \emph{comprehensive EEG understanding}. Existing evaluations, however, primarily target isolated decoding tasks or system-specific demonstrations, leaving the competence of large language models (LLMs) insufficiently quantified. We introduce \benchmarkname{}, a unified benchmark for comprehensive, instruction-conditioned EEG understanding. It comprises four subsets---Foundational Analysis, Sleep Assessment, Neurocognitive Assessment, and Physiological Integration---covering 17 datasets, \numcases{} tasks, and over \numinstances{} real-data instances. Given an instruction and EEG recordings with optional physiological signals, a system must perform the analysis and produce a scientifically grounded report and, when required, artifacts. Outputs are assessed through numerical, categorical, set, sequence, semantic, and artifact validation. We evaluate \nummodels{} representative LLMs across more than 100K executions under two paradigms: autonomous code execution with CodeAct and structured agentic analysis with BrainAgent. Results vary substantially across models, subsets, difficulty levels, and execution paradigms, showing that EEG competence depends on the model and its operationalization. \benchmarkname{} provides a reproducible testbed for advancing LLM-based EEG understanding. The code and benchmark will be released soon, with evaluation results continuously updated.

脑电分析大模型多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。