arXiv:2603.02268cs.LGcs.AI2026-03

用多样化脑电数据训练模型,提升临床诊断泛化能力。

PRISM: Exploring Heterogeneous Pretrained EEG Foundation Model Transfer to Clinical Differential Diagnosis

  • 构建跨人群、多设备的脑电预训练模型PRISM,对比窄源与广源数据效果
  • 在癫痫与伪差鉴别任务中,广源模型准确率高出12.3个百分点
  • 揭示评估基准差异会扭曲模型排名,提醒研究者注意数据处理细节

EEG基础模型通常在狭窄来源的临床数据集上预训练,并在同源基准上评估,难以判断其表征是源于神经生理还是记录分布偏差。本文提出PRISM(Population Representative Invariant Signal Model),一种沿预训练人群和下游适配双轴进行消融的掩码自编码器,架构与预处理保持一致。对比仅限欧盟/美国的窄源数据集(TUH + PhysioNet)与加入多中心南亚临床记录的广源数据集,三个发现浮现:第一,窄源预训练在匹配分布的基准上表现更强,而广源预训练在微调时更具适应性——此权衡在单一协议评估下不可见;在三个源数据集上训练的PRISM,在多数任务上达到或超越REVE(92个数据集,60,000+小时)水平,表明针对性多样性可替代盲目扩大规模,数据集数量是模型比较中的混淆变量。第二,在一项临床挑战性且此前未测试的任务——通过发作间期脑电图区分癫痫与诊断类比症中,广源检查点比窄源检查点高出12.3个百分点的平衡准确率,为所有评估中最大差距。第三,EEG-Bench与EEG-FM-Bench在相同数据集上对模型排名反向差异达24个百分点;我们识别出六种具体因素,包括数据划分方式、检查点选择、片段长度和归一化等,显示这些因素非加性地叠加影响结果。

原文摘要 · Abstract (English)

EEG foundation models are typically pretrained on narrow-source clinical archives and evaluated on benchmarks from the same ecosystem, leaving unclear whether representations encode neural physiology or recording-distribution artifacts. We introduce PRISM (Population Representative Invariant Signal Model), a masked autoencoder ablated along two axes -- pretraining population and downstream adaptation -- with architecture and preprocessing fixed. We compare a narrow-source EU/US corpus (TUH + PhysioNet) against a geographically diverse pool augmented with multi-center South Asian clinical recordings across multiple EEG systems. Three findings emerge. First, narrow-source pretraining yields stronger linear probes on distribution-matched benchmarks, while diverse pretraining produces more adaptable representations under fine-tuning -- a trade-off invisible under single-protocol evaluation. Trained on three source corpora, PRISM matches or outperforms REVE (92 datasets, 60,000+ hours) on the majority of tasks, demonstrating that targeted diversity can substitute for indiscriminate scale and that dataset count is a confounding variable in model comparison. Second, on a clinically challenging and previously untested task -- distinguishing epilepsy from diagnostic mimickers via interictal EEG -- the diverse checkpoint outperforms the narrow-source checkpoint by +12.3 pp balanced accuracy, the largest gap across all evaluations. Third, systematic inconsistencies between EEG-Bench and EEG-FM-Bench reverse model rankings on identical datasets by up to 24 pp; we identify six concrete sources including split construction, checkpoint selection, segment length, and normalization, showing these factors compound non-additively.

脑电分析模型泛化临床诊断预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。