用LRP解析脑电大模型,发现模型依赖眼动信号而非真实意图
From Clever Hans to Scientific Discovery: Interpreting EEG Foundational Transformers with LRP

- 用注意力感知的LRP方法分析脑电基础模型的决策依据
- 发现运动想象任务中模型误把眼动信号当动作信号
- 揭示情绪预测中对特定电极簇的反复依赖,或为唤醒的生理标志
脑电图(EEG)领域新兴的基础模型(FMs)在数据稀缺情况下仍有望推动深度学习在诊断与脑机接口中的应用,但其黑箱特性阻碍了广泛采纳。本文研究了注意力感知的逐层重要性传播(LRP)作为后处理归因方法,将传统用于卷积神经网络(CNN)EEG模型的LRP扩展至当前基础模型所依赖的Transformer架构。结果表明,LRP不仅能验证EEG-FM的决策合理性,还能从中揭示新的、具有生物学意义的假设:在运动想象任务中,模型表现出类似‘聪明汉斯’的行为,即优先关注与任务相关的眼动信号而非真正的运动相关信号;在自然情境下的情绪预测任务中,模型反复依赖中央电极集群,提示其可能构成唤醒的候选神经标记。尽管热图解释在此复杂领域仍存歧义,但本研究确立了LRP作为验证与探索EEG-FMs的工具,随着模型演进,其价值与发现潜力将持续提升。
原文摘要 · Abstract (English)
Emerging foundation models (FMs) in electroencephalography (EEG) promise a path to scale deep learning in diagnostics and brain-computer interfaces despite data scarcity, yet their opaque nature remains a barrier to wider adoption. We investigate attention-aware Layer-wise relevance propagation (LRP) as a post-hoc attribution method for EEG-FMs, extending LRP's use on convolutional neural network (CNN)-based EEG models to the Transformer architectures that current FMs are based on. We find that LRP can both verify EEG-FM decisions and surface novel, biologically plausible hypotheses from them. In motor imagery, it unmasks 'Clever Hans' behavior where models prioritize task correlated ocular signals over the intended motor correlates. In a naturalistic paradigm for affect prediction, it reveals a recurring reliance on a central electrode cluster, suggesting a candidate sensorimotor signature of arousal. Though heatmap interpretation remains ambiguous in this complex domain, the results position LRP as a tool for both verification and exploration of EEG-FMs, a role that will grow in both importance and discovery potential as the underlying models mature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。