arXiv:2512.19097cs.LGcs.AI2025-12被引 3

构建可适配多种脑电记录的通用模型,提升神经信号跨患者迁移能力。

DIVER-1: Scaling Intracranial EEG Foundation Models for Transferable Representations

  • 设计自适应注意力与时空重采样机制,兼容不同电极布局和采样率。
  • 在5310小时数据上预训练,性能超越现有模型,在癫痫检测与认知解码任务中领先。
  • 首次系统研究算力分配对脑电预训练的影响,发现数据量和训练时长更关键。

颅内脑电(iEEG)能直接获取毫秒级的人类神经活动记录,但由于电极布局、解剖覆盖范围、参考方式及记录条件在不同患者和中心间差异显著,可复用的表征学习面临挑战。本文提出DIVER-1,一种面向可变输入的自监督颅内脑电基础模型,结合任意电极-时间注意力、时空重采样、输入条件位置嵌入及多领域掩码重建,无需假设固定电极排列。我们在涵盖352,000通道小时的5,310小时ECoG与SEEG数据上预训练了两个版本:DIVER-1-0.1s和DIVER-1-1s,预训练总量约为BrainTreeBank的54倍。在两个独立测试集上评估:Neuroprobe用于自然情境认知解码,MAYO用于癫痫发作检测。在无泄漏信息的Neuroprobe上,尽管未使用源自Neuroprobe的BrainTreeBank数据进行预训练,DIVER-1-0.1s仍优于先前评估的iEEG基础模型,其平均AUROC超过线性光谱解码器,并保持与更强非线性基线相当的性能水平,此前的iEEG基础模型未能达到此高度。DIVER-1-1s在MAYO癫痫检测任务中也取得最高AUROC。最后,我们开展了首个受控的算力感知规模研究,系统探索数据量、受试者数量、训练时长和模型规模(最大达18亿参数)的缩放效应。结果表明存在数据受限区:增加独特记录数量并充分训练,比单纯扩大参数量更有效。代码已公开。

原文摘要 · Abstract (English)

Intracranial EEG (iEEG) provides direct, millisecond-scale recordings of human neural activity, but reusable representation learning is difficult because electrode layouts, anatomical coverage, referencing schemes, and recording conditions vary across patients and centers. We introduce DIVER-1, a self-supervised iEEG foundation model for variable-input recordings that combines any-variate electrode-time attention, spatio-temporal resampling, input-conditioned positional embeddings, and multi-domain masked reconstruction without assuming a fixed electrode montage. We pretrain two variants, DIVER-1-0.1s and DIVER-1-1s, on 5,310 hours of ECoG and SEEG spanning 352k channel-hours, roughly 54x the BrainTreeBank-based pretraining volume. We evaluate DIVER-1 on two held-out benchmarks: Neuroprobe for naturalistic cognitive decoding and MAYO for seizure detection. On leakage-aware Neuroprobe, DIVER-1-0.1s outperforms prior evaluated iEEG foundation models despite using no BrainTreeBank recordings, the corpus underlying Neuroprobe, during pretraining; it also exceeds the linear spectrogram decoder in mean AUROC and remains competitive with stronger nonlinear baselines, a level prior evaluated iEEG foundation models did not reach. DIVER-1-1s also achieves the top AUROC on MAYO seizure detection. Finally, we conduct, to our knowledge, the first controlled compute-aware scaling study for self-supervised iEEG pretraining, sweeping data scale, subject count, training duration, and model size up to 1.8B parameters. Our results indicate a data-constrained regime: expanding unique recordings and training sufficiently long are more reliable scaling axes than increasing parameter count alone. Code is available at link.

脑电建模自监督学习跨患者迁移神经信号分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。