让脑电与语言对齐,提升脑机接口的解码能力
LEAF: Language-EEG Aligned Foundation Model for Brain-Computer Interfaces
- 用语言指令指导脑电信号学习,生成语义一致的嵌入
- 在16个数据集上12项领先,五类任务平均表现最优
- 首次证明语言指令能引导脑电空间走向语义连贯
近年来,脑电图(EEG)基础模型在捕捉可迁移的脑电表征方面取得进展,显著加速了脑机接口(BCI)的发展。然而,现有方法难以将语言指令作为先验约束融入脑电表征学习,限制了利用语言内在语义知识统一不同标签和任务的能力。为此,我们提出LEAF——一种面向脑电-语言对齐的基础模型,支持语义任务指令与查询。该模型引入任务感知的语义引导,生成结构化且语言对齐的脑电嵌入,从而提升解码鲁棒性与可迁移性。预训练阶段,我们设计联合频谱-时序重建(STR)框架,捕捉脑电信号的耦合频谱节律与时序动态;通过随机频谱扰动增强频率鲁棒性,并采用两个互补的时序目标学习上下文与序列结构。在脑电-语言对齐阶段,提出指令条件的Q-Former(IQF),基于查询的交叉注意力变压器将指令嵌入注入脑电令牌,通过可学习查询实现与文本标签嵌入的语义对齐。我们在16个下游数据集上评估LEAF,涵盖运动想象、情绪识别、稳态视觉诱发电位、隐含言语及医疗任务。结果表明,LEAF在12个数据集上达到最先进水平,在五个任务类别中平均表现最佳。重要的是,分析首次揭示显式任务指令作为语义先验,能够引导脑电嵌入进入连贯且语言基础的空间。代码与预训练权重将公开发布。
原文摘要 · Abstract (English)
Recent advances in electroencephalography (EEG) foundation models, which capture transferable EEG representations, have greatly accelerated the development of brain-computer interfaces (BCIs). However, existing approaches still struggle to incorporate language instructions as prior constraints for EEG representation learning, limiting their ability to leverage the semantic knowledge inherent in language to unify different labels and tasks. To address this challenge, we present LEAF, a foundation model for EEG--Language Alignment with Semantic Task Instruction and Querying. LEAF integrates task-aware semantic guidance to produce structured and linguistically aligned EEG embeddings, thereby enhancing decoding robustness and transferability. In the pretraining stage, we introduce a joint Spectral--Temporal Reconstruction (STR) framework that captures the coupled spectral rhythms and temporal dynamics of EEG signals. STR applies randomized spectral perturbation to enhance frequency robustness and uses two complementary temporal objectives to learn both contextual and sequential structure. In the EEG-Language alignment stage, we propose the Instruction-conditioned Q-Former (IQF). This query-based cross-attention transformer injects instruction embeddings into EEG tokens and achieves semantic alignment with textual label embeddings through learnable queries. We evaluate LEAF on 16 downstream datasets spanning motor imagery, emotion recognition, steady-state visual evoked potentials, covert speech, and healthcare tasks. LEAF achieves state-of-the-art performance on 12 of the 16 datasets and obtains the best average results across all five task categories. Importantly, our analyses reveal for the first time that explicit task instructions serve as semantic priors guiding EEG embeddings into coherent and linguistically grounded spaces. The code and pre-trained weights will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。