arXiv:2606.22925cs.LGcs.NE2026-06

为脑电图研究建立可执行的任务规范层,让不同数据集能统一评估。

EEG Benchmarking Needs a Task Specification Layer: NeuroDoc for Rulebook-Guided, Executable Benchmark Construction

论文配图:EEG Benchmarking Needs a Task Specification Layer: NeuroDoc for Rulebook-Guided, Executable Benchmark Construction
图 1 · 摘自论文原文
  • 用任务文档+可执行内核构建标准化评测单元
  • 发布含245个任务定义的53项已审核基准条目
  • 适合脑电图模型开发者与标准制定者使用

脑电图(EEG)基础模型越来越依赖多数据集训练与评估,但公开数据集仍缺乏统一的任务规范层,无法将异构记录转化为可复用的基准单元。现有标准仅组织文件、元数据和来源信息,未以通用语言和规则书定义任务,导致关键任务语义分散于论文、代码和人工解读中。本文探究通过结构化任务规范语言与共享规则书,对异构公开EEG数据集进行标准化的可行性。方法将每个基准条目表示为与可执行任务内核同步的任务文档,规则书定义任务字段、证据要求、文档-内核对齐、评审状态及机器可验证约束。基于此,我们发布一个由社区评审的EEG基准语料库,包含53个已完成且经审核的条目,涵盖245个任务定义,覆盖多种实验范式,并引入NeuroDoc与NeuroAudit作为规则书引导的起草、升级、评审、修改与发布管理的操作支持层。进一步在四个EEG基础模型骨干网络上测试这些基准单元的跨模型可实例化能力,提供可执行、可审计、可复用的脑电图基准基础设施的实证支持。

原文摘要 · Abstract (English)

Electroencephalography (EEG) foundation models increasingly rely on multi-dataset training and evaluation, yet public EEG datasets still lack a shared task specification layer that can turn heterogeneous recordings into reusable benchmark units. Existing standards organize files, metadata, and provenance, but they do not specify EEG tasks under a common language and rulebook, leaving critical task semantics scattered across papers, code, and manual interpretation. We investigate whether heterogeneous public EEG datasets can be standardized through a structured task specification language paired with a shared rulebook. Our methodology represents each benchmark entry as a task document synchronized with an executable task kernel, with the rulebook defining task fields, evidence requirements, document-kernel alignment, review states, and machine-checkable constraints. Using this methodology, we release a community-reviewed EEG benchmark corpus centered on 53 completed and reviewed entries with 245 task definitions spanning diverse paradigms, and we introduce NeuroDoc and NeuroAudit as the operational support layer for rulebook-guided drafting, upgrading, review, amendment, and release management. We further examine whether the resulting benchmark units can be instantiated in a shared downstream setting across four EEG foundation model backbones, providing execution-based evidence for reusable, auditable, and executable EEG benchmarking infrastructure.

脑电图基准评测可执行规则书

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。