让脑电波理解对话,跨模态对齐提升神经信号解析
WaveMind: Towards a Conversational EEG Foundation Model Aligned to Textual and Visual Modalities
- 将脑电信号与文本/视觉信息映射到统一语义空间
- 在4个下游任务中实现高精度分类与自由对话能力
- 首个用于指令微调的跨任务脑电数据集,支持通用建模
利用多模态大语言模型分析脑电图(EEG)为解读大脑活动提供了新路径。然而,脑电信号同时包含认知过程与内在神经状态,导致配对数据模态不匹配,阻碍跨模态表征学习。通过深入研究,我们发现这些模态间存在互补关系。基于此,提出将脑电信号及其对应模态映射至统一语义空间,实现泛化解读。为进一步支持对话功能,构建了首个用于指令微调的跨任务脑电数据集WaveMind-Instruct-338k。所获模型在4个下游任务中均表现出稳健分类性能,并支持灵活开放的对话,为神经科学研究及通用脑电模型开发提供重要参考。
原文摘要 · Abstract (English)
Electroencephalography (EEG) interpretation using multimodal large language models (MLLMs) offers a novel approach for analyzing brain signals. However, the complex nature of brain activity introduces critical challenges: EEG signals simultaneously encode both cognitive processes and intrinsic neural states, creating a mismatch in EEG paired-data modality that hinders effective cross-modal representation learning. Through a pivot investigation, we uncover complementary relationships between these modalities. Leveraging this insight, we propose mapping EEG signals and their corresponding modalities into a unified semantic space to achieve generalized interpretation. To fully enable conversational capabilities, we further introduce WaveMind-Instruct-338k, the first cross-task EEG dataset for instruction tuning. The resulting model demonstrates robust classification accuracy while supporting flexible, open-ended conversations across four downstream tasks, thereby offering valuable insights for both neuroscience research and the development of general-purpose EEG models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。