arXiv:2605.13930cs.LGcs.HC2026-05被引 3

用稀疏自编码器解析脑电大模型内部机制,揭示临床概念混杂与干预失效问题。

Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders

论文配图:Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders
图 1 · 摘自论文原文
  • 通过TopK稀疏自编码器提取三类脑电变换器的稀疏特征
  • 发现年龄与病理混杂等关键表征缺陷,导致干预时性能崩溃
  • 提出频谱解码器将隐空间操作映射为可解释的脑电波段变化

脑电基础模型在临床任务中表现卓越,但其内部决策过程仍不透明,制约临床信任。本文对三种架构不同的脑电变压器(SleepFM、REVE、LaBraM)应用TopK稀疏自编码器(SAEs),从中提取稀疏特征字典,并基于异常、年龄、性别、药物等临床分类体系评估单义性与纠缠度。仅通过一个超参数配置,即可在三类模型间稳健迁移。通过概念操控实验,引入“目标与非目标”探针区域度量以量化操控选择性,识别出三种操作模式:可选择性操控、编码但纠缠、未编码。该框架揭示了重大表征缺陷:‘破坏性’干预会全局降低模型性能,以及年龄-病理混淆等临床纠缠现象——无法独立抑制某一概念而不损害另一概念。最后,通过谱解码器将这些干预映射回振幅谱,实现对潜在生理信号的可解释性转化,如病理慢波抑制和α波恢复。

原文摘要 · Abstract (English)

EEG foundation models achieve state-of-the-art clinical performance, yet the internal computations driving their predictions remain opaque: a barrier to clinical trust. We apply TopK Sparse Autoencoders (SAEs) across three architecturally distinct EEG transformers: SleepFM, REVE, and LaBraM to extract sparse feature dictionaries from their embeddings. By grounding these features in a clinical taxonomy (abnormality, age, sex, and medication), we benchmark monosemanticity and entanglement across architectures. A single hyperparameter procedure, driven by an intrinsic dictionary health audit, transfers robustly across all three architectures. Via concept steering, we introduce a "target vs. off-target" probe area metric to quantify steering selectivity and reveal three operational regimes: selectively steerable, encoded but entangled, and non-encoded. This framework exposes critical representational failures: "wrecking-ball" interventions that collapse global model performance, and clinical entanglements, such as age-pathology confounding, where it is impossible to suppress one concept without corrupting the other. Finally, a spectral decoder maps these interventions back to the amplitude spectrum, translating latent manipulations into physiologically interpretable frequency signatures, such as pathological slow-wave suppression and $α$-band restoration.

脑电分析可解释性稀疏编码神经科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。