arXiv:2606.06647cs.LGq-bio.NC2026-06被引 2

发现脑电基础模型存在身份陷阱,提出可诊断该问题的分析框架。

The Identity Trap in EEG Foundation Models: A Diagnostic Audit

论文配图:The Identity Trap in EEG Foundation Models: A Diagnostic Audit
图 1 · 摘自论文原文
  • 设计冻结表征诊断协议,通过五种方法检测身份线索。
  • 89倍于随机基线的身份方差在12组实验中普遍存在,且微调后增强。
  • 适用于研究脑电模型是否真正学习生物标记而非个体特征的研究者。

EEG基础模型在临床静息态脑电上报告高准确率,但跨被试交叉验证下的高准确率可能源于真实生物标记,或与被试身份相关的特征。我们称之为身份陷阱,并探讨能否在微调前通过表征层面诊断。提出FMScope:一种冻结表征协议,包含方差分解、被试轴擦除、非周期性1/f扰动、逐层标签探测和同被试方向一致性等五项诊断。在三个预训练模型(LaBraM、CBraMod、REVE)及四个数据集上进行2×2布局测试:标签关联性 × 是否存在共识跨被试脑电标记。结果表明:(i) 身份陷阱普遍存在:冻结状态下被试方差为随机基线的13-89倍,在全部12组中上升(微调后+10至+63个百分点);移除该线性轴可提升标签解码性能(主要细胞+6至+12个百分点,外部队列+4至+27个百分点)。(ii) 非周期性1/f是身份载体之一:去除后,LaBraM与CBraMod的被试探测下降9-19个百分点;而REVE无明显依赖。(iii) 微调仅在具有文献支持的跨被试标记的脑区增强标签方差。结论:身份陷阱是物理可解释的捷径学习现象,其首选线索具有生理基础,仅靠被试级划分无法排除。FMScope能区分生物学标记与身份信号带来的性能提升。

原文摘要 · Abstract (English)

Objective. EEG foundation models (FMs) report strong accuracy on clinical resting-state EEG. However, high accuracy under subject-disjoint cross-validation remains ambiguous: it can reflect a genuine clinical biomarker, or subject-identity features that correlate with the label. We name this the Identity Trap and ask whether it can be diagnosed at the representation level before fine-tuning. Approach. We propose FMScope, a frozen-representation protocol packaging five diagnostics: variance decomposition, subject-axis erasure, aperiodic 1/f ablation, layer-wise label probing, and within-subject direction consistency. We apply it to three pretrained FMs (LaBraM, CBraMod, REVE) across four datasets in a 2x2 layout: subject relation of label x presence of a consensus cross-subject EEG marker. Main results. (i) The Identity Trap is universal: frozen subject-variance is 13-89x a random null in 12/12 pairs, rising in all 12 under fine-tuning (+10 to +63 pp). This dominance is a removable linear axis: erasing it improves label decoding where the label varies within subject (+6 to +12 pp in primary cells; +4 to +27 pp across external cohorts). (ii) Aperiodic 1/f is one subject carrier: removing it drops the subject probe by 9-19 pp on LaBraM and CBraMod. REVE saturates subject identity without measurable aperiodic dependence. (iii) Fine-tuning amplifies label-variance only in cells with a literature-established cross-subject marker. Significance. The Identity Trap is a physically-grounded instance of shortcut learning: the preferred cue has a measurable physiological component, and subject-disjoint splitting alone cannot rule it out. FMScope separates gains reflecting a biological marker from those reflecting subject identity.

脑电模型诊断身份陷阱基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。