arXiv:2605.11410cs.AI2026-05被引 4

揭秘脑电基础模型到底学到了什么,发现其核心能力可被已有特征解释七成以上。

What Do EEG Foundation Models Capture from Human Brain Signals?

论文配图:What Do EEG Foundation Models Capture from Human Brain Signals?
图 1 · 摘自论文原文
  • 用层间探针和消融实验分析模型学到的脑电信号表征机制。
  • 68.6%的特征与任务表现直接相关,50个特征在多任务中具普适性。
  • 模型优势近八成可由现有脑电特征解释,难任务仍留出改进空间。

临床脑电图(EEG)分析长期依赖人工设计的特征体系,如频带功率、连通性、复杂度等。现代脑电基础模型通过自监督预训练直接从原始信号学习,多数任务上表现超越传统特征工程方法。但二者表征是否对齐仍是未解之谜。本文从‘模型学了什么’‘用了什么’‘能解释多少’三方面展开评估,采用层间岭探针、跨协方差子空间消融及透明分类器基准测试。涵盖三个基础模型(CSBrain, CBraMod, LaBraM)、五项临床任务(MDD, Stress, ISRUC-Sleep, TUSL, Siena)及6类共63个特征。在945个(模型,任务,特征)组合中,648个(68.6%)为因果相关,199个(21.1%)仅为编码存在。50个特征在两个以上任务中被三架构一致支持,属潜在通用候选。频域特征主导,其余五类亦贡献显著因果质量。经验证,这些特征平均恢复模型相比随机基线79.3%的优势,任务间呈现清晰梯度:重度抑郁(MDD)接近完全恢复(≈0.99),压力检测则仅恢复至≈0.56,表明硬任务仍有明确概念探索方向。

原文摘要 · Abstract (English)

Clinical electroencephalogram (EEG) analysis rests on a hand-crafted feature catalog refined over decades, \emph{e.g.,} band power, connectivity, complexity, and more. Modern EEG foundation models bypass this catalog, learn directly from raw signals via self-supervised pretraining, and match or outperform feature-engineered baselines on most clinical benchmarks. Whether the two representations align is an open question, which we decompose into three sub-questions: \emph{what does the model learn}, \emph{what does the model use}, and \emph{how much can be explained}. We answer them with layer-wise ridge probing, LEACE-style cross-covariance subspace erasure, and a transparent classifier benchmarked against a random-feature baseline. The audit covers three foundation models (CSBrain, CBraMod, LaBraM), five clinical tasks (MDD, Stress, ISRUC-Sleep, TUSL, Siena), and a 6-family 63-feature lexicon. Of the $945$ (model, task, feature) units, $648$ ($68.6\%$) are representation-causal and $199$ ($21.1\%$) are encoded-only. Across tasks, $50$ features qualify as universal candidates with strong support (all three architectures RC) in two or more tasks. Frequency-domain features dominate, but the other five families each contribute substantial causal mass. Confirmed features recover, on average, $79.3\%$ of the foundation model's advantage over the random baseline, with a clean task gradient (MDD $\approx 0.99$ down to Stress $\approx 0.56$): tasks near ceiling are almost fully recovered by the lexicon, while harder tasks leave a non-trivial residual that pinpoints a concrete target for future concept discovery.

脑电分析基础模型可解释性特征挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。