构建多模态脑电基础模型,实现跨数据集的癫痫检测泛化。
Multimodal Pretraining for Generalizable EEG Representation Learning

- 融合原始信号、时频图与文本的多模态编码器,共享嵌入空间。
- 无标签预训练下达到CHB-MIT数据集0.874 AUROC,LOSO平均准确率0.558。
- 支持跨患者泛化与可解释定位,适合临床部署与新场景快速适配。
用于癫痫的脑电(EEG)模型通常局限于特定数据集和任务,难以跨数据集或情境应用。近期基础模型与自监督学习研究提示,可构建适应性强的EEG主干网络。本文提出一种多模态EEG基础模型,包含基于Mamba架构的原始信号编码器、类视觉变压器(ViT)的时频数据编码器,以及轻量级文本编码器,三者共享嵌入空间。预训练采用掩码建模、跨视图对比对齐与时间一致性损失等创新技术,无需标注数据即可生成富含癫痫相关性的表示。在标准CHB-MIT癫痫检测基准上,单模型最优达0.874 AUROC,集成模型达0.878,为当前最佳性能。进一步在罕见报告的留一被试排除(LOSO)协议下评估,19名被试平均平衡准确率为0.558,凸显患者无关检测难度。该模型在多数据集与评估设置中均表现稳健,支持癫痫检测的鲁棒实现与新场景快速迁移,并具备可解释性癫痫定位能力。
原文摘要 · Abstract (English)
Electroencephalography (EEG) models used for epilepsy are often limited to specific datasets and tasks. This limited approach can make it challenging to apply these models across different datasets or in various situations. However, recent studies in foundation models and self-supervised learning suggest that an adaptable EEG backbone could support a range of EEG related tasks. In this study, we have developed a multimodal EEG foundation model that combines a raw signal encoder based on the Mamba architecture, a Vision Transformer (ViT)-style encoder for time-frequency data, and a lightweight encoder for text, all within a shared embedding space. The pretraining process relies on several innovative techniques, such as masked modeling, cross-view contrastive alignment, and temporal consistency losses. These methods are designed to create rich, seizure-relevant representations without requiring labeled data. To assess the efficacy and generalization of our pretrained model, we fine-tuned it on the canonical CHB-MIT seizure detection benchmark and additional seizure detection datasets, and conducted extensive experiments comparing different model variants. On the standard CHB-MIT split, our best single model achieved an AUROC of 0.874, and an ensemble variant reached 0.878 AUROC, representing state-of-the-art performance on this benchmark. In addition to standard train-test splits, we evaluated performance under a leave-one-subject-out (LOSO) protocol, which is rarely reported in prior EEG seizure modeling work and highlights the difficulty of patient-independent seizure detection, with a mean LOSO balanced accuracy of 0.558 across 19 subjects. Across datasets and evaluation settings, our multimodal foundation model enabled robust seizure detection and straightforward adaptation to new seizure detection scenarios, while also supporting interpretable seizure localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。