arXiv:2605.10817cs.AI2026-05被引 1

构建可理解临床背景的长时脑电基础模型,提升疾病诊断准确率。

CLEF: EEG Foundation Model for Learning Clinical Semantics

论文配图:CLEF: EEG Foundation Model for Learning Clinical Semantics
图 1 · 摘自论文原文
  • 用三维谱图令牌表示整段脑电数据,支持大尺度建模。
  • 在234项任务中平均AUC提升至0.74,优于以往模型229项。
  • 融合医生报告与电子病历,适合临床研究与医疗AI开发者使用。

临床脑电图解读需综合分析完整脑电记录并结合临床背景。现有脑电基础模型多聚焦短窗解码,缺乏临床上下文整合。本文提出CLEF,一种面向临床场景的长时脑电基础模型。CLEF将脑电会话表示为3D多窗谱图令牌,实现会话级的可处理Transformer建模,并通过对比学习对齐神经科医生报告与结构化电子病历(EHR)数据。在包含234个任务的新基准上评估,涵盖疾病表型、用药暴露和脑电发现,数据来自超过108,000名患者,共26万余段脑电记录。CLEF在234项任务中优于先前模型229项,平均AUROC从0.65提升至0.74。仅重建预训练已超越以往模型,结合报告与EHR对齐进一步提升性能。留出概念与外部队列实验表明,该表示具有跨任务迁移能力。结果表明,以临床为导向的会话级表示学习是临床脑电基础模型的可行范式。

原文摘要 · Abstract (English)

Clinical EEG interpretation requires reasoning over full EEG sessions and integrating signal patterns with clinical context. Existing EEG foundation models are largely designed for short-window decoding and do not incorporate clinical context. We introduce CLEF, a clinically grounded long-context EEG foundation model. CLEF represents EEG sessions as 3D multitaper spectrogram tokens, enabling tractable Transformer modeling at session scale, and aligns embeddings with neurologist reports and structured EHR data through contrastive objectives. We evaluate CLEF on a new 234-task benchmark spanning disease phenotypes, medication exposures, and EEG findings, with more than 260k EEG sessions from over 108k patients. CLEF outperforms prior EEG foundation models on 229 of 234 tasks, improving mean AUROC from 0.65 to 0.74. Reconstruction-only pretraining surpasses prior EEG foundation models, while report and EHR alignment yields further gains. Held-out concept and external-cohort experiments suggest that these representations transfer beyond observed alignment targets. These results support session-scale, clinically grounded representation learning as a promising foundation-model paradigm for clinical EEG.

脑电图基础模型临床医学长序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。